Home › Whisper accuracy by language

Benchmark

Whisper accuracy across 29 languages

Most published Whisper benchmarks measure English. This one measures 29 languages with six model variants on Apple Silicon, using the same decoder settings a real app ships with — and it finds that the 547 MB default holds in 23 of them, while the model directly above it is worse in 9.

100%on-device — audio never uploaded
Freeto download — Pro unlock $49.99 once
99languages, auto-detected
macOS 13+Apple Silicon (M1+)

What was measured

Six Whisper model variants transcribed the same read-speech corpus in 29 languages — every language the Lesskeys App Store listing ships in — on Apple Silicon, through the same decoder settings the app itself uses (greedy, no beam search). Rates are word error rate, except for Japanese, Korean, Thai and Chinese, where words are not space-delimited and character error rate is used instead. Lower is better.

The results

LanguageMetrictinysmallturbo-q5turbov3-q5v3
Arabic ar-SAWER67.0%36.9%22.9%22.4%22.3%22.2%
Chinese (Simplified) zh-HansCER55.0%24.8%13.7%16.8%13.1%12.6%
Czech csWER84.0%40.4%13.2%12.8%13.1%11.9%
Danish daWER87.2%35.0%13.4%13.0%12.6% 13.3%
Dutch nl-NLWER74.8%28.8%10.8%9.8%8.3% 8.4%
English en-USWER31.9%10.7%9.1%8.9%7.0% 8.9%
Finnish fiWER64.4%31.8%10.8%10.1%7.7% 7.7%
French fr-FRWER50.1%26.1%10.0%10.1%7.7%7.5%
German de-DEWER35.0%11.9%5.2%4.7% 5.2%5.2%
Greek elWER74.7%32.9%11.2% 11.3%11.8%12.3%
Hebrew heWER81.9%63.2%51.0%48.4%48.3% 48.3%
Hindi hiWER100.0%68.3%34.4%42.1%33.4%32.4%
Hungarian huWER84.8%48.7%17.9%17.9%17.0%16.3%
Indonesian idWER59.2%21.8%7.4% 7.8%8.1%9.4%
Italian itWER31.8%8.2%4.0%2.9% 4.6%4.5%
Japanese jaCER54.3%25.4%18.5%15.6%14.2% 17.0%
Korean koCER33.9%24.6%11.3%11.3%9.4% 9.5%
Malay msWER62.5%25.1%7.1%9.8%6.0%5.9%
Norwegian noWER64.0%24.2%10.0%9.8% 10.0%9.8%
Polish plWER57.0%17.7%6.7%6.4% 7.4%7.6%
Portuguese (BR) pt-BRWER25.4%7.9%5.2%5.1%5.0%4.9%
Romanian roWER79.1%33.4%10.2%9.9%9.5%9.2%
Russian ruWER38.7%12.9%6.4%6.3% 6.8%7.3%
Spanish es-ESWER19.3%7.0%3.5%3.5%3.0%2.8%
Swedish svWER57.2%21.4%8.3%8.1%7.1% 7.3%
Thai thCER62.4%27.5%20.8%24.6%13.2%12.6%
Turkish trWER62.1%30.5%15.2%15.5%13.1% 13.5%
Ukrainian ukWER68.3%39.2%23.1%21.0% 23.3%23.1%
Vietnamese viWER64.1%26.6%9.8% 10.3%10.2%10.2%

● marks the best model for that language. The highlighted column is the 547 MB model Lesskeys downloads by default.

Four things the data says

1. The 547 MB default holds in 23 of 29 languages

In 23 languages, no model in the set is significantly better than the default — including large-v3 at 3.0 GB, more than five times the download. "Significantly" here means a two-proportion z-test at z ≤ −2 against the default; at this sample size a difference that fails that test is not distinguishable from noise, so read those 23 as not separated rather than proven equal. If you are choosing a model to run locally, the small quantized one is very often the right answer.

2. 6 languages where a bigger model genuinely earns its size

LanguageDefault (547 MB)Better modelGain
Thai20.8%v3 — 12.6%−8.2 pp
Japanese18.5%v3-q5 — 14.2%−4.3 pp
Finnish10.8%v3-q5 — 7.7%−3.1 pp
French10.0%v3 — 7.5%−2.5 pp
Dutch10.8%v3-q5 — 8.3%−2.5 pp
Korean11.3%v3-q5 — 9.4%−1.9 pp

If you work in one of these, the larger download is worth it. Note that v3-q5 (1.0 GB) captures nearly all of the available gain: across all 29 languages the 3.0 GB v3 never improves on it by more than 1.2 percentage points. The extra 2 GB buys very little.

3. Bigger is not monotonically better

turbo (1.6 GB) — the model directly above the default in most pickers, at three times the download — scores numerically worse than the 547 MB default in 9 of 29 languages: Hindi (42.1% vs 34.4%), Thai (24.6% vs 20.8%), Chinese (Simplified) (16.8% vs 13.7%), Malay (9.8% vs 7.1%), Vietnamese (10.3% vs 9.8%), Indonesian (7.8% vs 7.4%), Turkish (15.5% vs 15.2%), Greek (11.3% vs 11.2%), French (10.1% vs 10.0%). (Margins that small are not all individually significant at this sample size; the point is the direction, which size alone would not predict.) Reaching for the next model up is not a safe default.

4. tiny is unusable in most of the world

The 75 MB model is frequently reported as a viable lightweight option. Measured across 29 languages it is not: Hindi 100.0%, Danish 87.2%, Hungarian 84.8%, Czech 84.0%, Hebrew 81.9%. It is a placeholder that loads while a real model downloads, not a model to work in.

The hardest and easiest languages

The best result in the entire sweep is Spanish at 2.8%. The worst is Hebrew, which does not fall below 48.3% in any configuration tested — no model choice fixes it. If your work is in that language, local Whisper transcription of any size will need correction.

Related: silent truncation — when a model stops early and the app reports success →

Correction: the Indic and Arabic rows are under re-measurement

Dated 2026-08-28. The scorer behind this table normalised text with [^\w\s], and Python's \w does not match Unicode combining marks. Every Devanagari vowel sign, Bengali nukta and Arabic diacritic was therefore replaced by a space, splitting one word into several. Measured on the same references this page used: Tamil scored 1,520 units against 437 real words, Hindi 1,237 against 683, Assamese 1,220 against 498. Latin, Cyrillic and CJK rows are unaffected — those scripts carry no combining marks — and so are the conclusions that rest on them.

This inflates the reference and, because fragments cannot align, the error count with it. So the rows for Hindi, Arabic, Hebrew and the other combining-mark scripts overstate the error rate by an amount that varies per language, and the “worst language” claim above is not safe for those rows. A corrected sweep across 83 languages is running now and this page will be regenerated from it.

It is recorded here rather than quietly deleted, for the same reason as the retraction below: a normaliser that silently mangles non-Latin script is exactly the failure a reader of a multilingual benchmark should be told about, and it is not visible in any published number.

A retraction, kept in place

An earlier version of this sweep reported large-v3 as significantly better than the default for Ukrainian. That was wrong. It was measured under beam 5, which is whisper-cli's default but not what the app ships. Re-run greedy — the real setting — the default and v3 land in a dead tie. The gap was the decoder, not the model. It is recorded here rather than quietly deleted, because the decoder setting is exactly the variable most local-Whisper comparisons leave unstated.

Scope and limits

Download on the Mac App Store

Free to download. One-time unlock. No subscription. · Pro $49.99 one-time · macOS 13+

FAQ

Which Whisper model should I run on a Mac?

In 23 of the 29 languages measured, the 547 MB quantized Large V3 Turbo model is as accurate as anything larger, so it is the sensible default. The exceptions are listed on this page; in those languages the 1.0 GB v3-q5 model is worth the extra download.

Does quantizing a Whisper model hurt accuracy?

Barely. Across all 29 languages the full-precision 3.0 GB model never beats the 1.0 GB quantized one by more than 1.2 percentage points, and the 547 MB quantized default matches much larger models in most languages.

Is a bigger Whisper model always more accurate?

No. The 1.6 GB turbo model scores numerically worse than the 547 MB default in 9 of the 29 languages measured. Model size is not a reliable proxy for accuracy in a given language.

Why is word error rate not used for Japanese, Korean, Thai and Chinese?

Those languages are not space-delimited, so word-level alignment is not meaningful. Character error rate is the standard substitute and is what this page reports for them.

Was this measured in the cloud?

No. Every number was produced locally on Apple Silicon, which is also how Lesskeys runs transcription — the audio never leaves the machine.

Related