Home › Whisper accuracy by language
Benchmark
Whisper accuracy across 29 languages
Most published Whisper benchmarks measure English. This one measures 29 languages with six model variants on Apple Silicon, using the same decoder settings a real app ships with — and it finds that the 547 MB default holds in 23 of them, while the model directly above it is worse in 9.
What was measured
Six Whisper model variants transcribed the same read-speech corpus in 29 languages — every language the Lesskeys App Store listing ships in — on Apple Silicon, through the same decoder settings the app itself uses (greedy, no beam search). Rates are word error rate, except for Japanese, Korean, Thai and Chinese, where words are not space-delimited and character error rate is used instead. Lower is better.
The results
| Language | Metric | tiny | small | turbo-q5 | turbo | v3-q5 | v3 |
|---|---|---|---|---|---|---|---|
| Arabic ar-SA | WER | 67.0% | 36.9% | 22.9% | 22.4% | 22.3% | 22.2% ● |
| Chinese (Simplified) zh-Hans | CER | 55.0% | 24.8% | 13.7% | 16.8% | 13.1% | 12.6% ● |
| Czech cs | WER | 84.0% | 40.4% | 13.2% | 12.8% | 13.1% | 11.9% ● |
| Danish da | WER | 87.2% | 35.0% | 13.4% | 13.0% | 12.6% ● | 13.3% |
| Dutch nl-NL | WER | 74.8% | 28.8% | 10.8% | 9.8% | 8.3% ● | 8.4% |
| English en-US | WER | 31.9% | 10.7% | 9.1% | 8.9% | 7.0% ● | 8.9% |
| Finnish fi | WER | 64.4% | 31.8% | 10.8% | 10.1% | 7.7% ● | 7.7% ● |
| French fr-FR | WER | 50.1% | 26.1% | 10.0% | 10.1% | 7.7% | 7.5% ● |
| German de-DE | WER | 35.0% | 11.9% | 5.2% | 4.7% ● | 5.2% | 5.2% |
| Greek el | WER | 74.7% | 32.9% | 11.2% ● | 11.3% | 11.8% | 12.3% |
| Hebrew he | WER | 81.9% | 63.2% | 51.0% | 48.4% | 48.3% ● | 48.3% ● |
| Hindi hi | WER | 100.0% | 68.3% | 34.4% | 42.1% | 33.4% | 32.4% ● |
| Hungarian hu | WER | 84.8% | 48.7% | 17.9% | 17.9% | 17.0% | 16.3% ● |
| Indonesian id | WER | 59.2% | 21.8% | 7.4% ● | 7.8% | 8.1% | 9.4% |
| Italian it | WER | 31.8% | 8.2% | 4.0% | 2.9% ● | 4.6% | 4.5% |
| Japanese ja | CER | 54.3% | 25.4% | 18.5% | 15.6% | 14.2% ● | 17.0% |
| Korean ko | CER | 33.9% | 24.6% | 11.3% | 11.3% | 9.4% ● | 9.5% |
| Malay ms | WER | 62.5% | 25.1% | 7.1% | 9.8% | 6.0% | 5.9% ● |
| Norwegian no | WER | 64.0% | 24.2% | 10.0% | 9.8% ● | 10.0% | 9.8% ● |
| Polish pl | WER | 57.0% | 17.7% | 6.7% | 6.4% ● | 7.4% | 7.6% |
| Portuguese (BR) pt-BR | WER | 25.4% | 7.9% | 5.2% | 5.1% | 5.0% | 4.9% ● |
| Romanian ro | WER | 79.1% | 33.4% | 10.2% | 9.9% | 9.5% | 9.2% ● |
| Russian ru | WER | 38.7% | 12.9% | 6.4% | 6.3% ● | 6.8% | 7.3% |
| Spanish es-ES | WER | 19.3% | 7.0% | 3.5% | 3.5% | 3.0% | 2.8% ● |
| Swedish sv | WER | 57.2% | 21.4% | 8.3% | 8.1% | 7.1% ● | 7.3% |
| Thai th | CER | 62.4% | 27.5% | 20.8% | 24.6% | 13.2% | 12.6% ● |
| Turkish tr | WER | 62.1% | 30.5% | 15.2% | 15.5% | 13.1% ● | 13.5% |
| Ukrainian uk | WER | 68.3% | 39.2% | 23.1% | 21.0% ● | 23.3% | 23.1% |
| Vietnamese vi | WER | 64.1% | 26.6% | 9.8% ● | 10.3% | 10.2% | 10.2% |
● marks the best model for that language. The highlighted column is the 547 MB model Lesskeys downloads by default.
Four things the data says
1. The 547 MB default holds in 23 of 29 languages
In 23 languages, no model in the set is significantly better than the default
— including large-v3 at 3.0 GB, more than five times the download.
"Significantly" here means a two-proportion z-test at z ≤ −2 against the default;
at this sample size a difference that fails that test is not distinguishable from noise, so read
those 23 as not separated rather than proven equal. If you are choosing a model to run
locally, the small quantized one is very often the right answer.
2. 6 languages where a bigger model genuinely earns its size
| Language | Default (547 MB) | Better model | Gain |
|---|---|---|---|
| Thai | 20.8% | v3 — 12.6% | −8.2 pp |
| Japanese | 18.5% | v3-q5 — 14.2% | −4.3 pp |
| Finnish | 10.8% | v3-q5 — 7.7% | −3.1 pp |
| French | 10.0% | v3 — 7.5% | −2.5 pp |
| Dutch | 10.8% | v3-q5 — 8.3% | −2.5 pp |
| Korean | 11.3% | v3-q5 — 9.4% | −1.9 pp |
If you work in one of these, the larger download is worth it. Note that
v3-q5 (1.0 GB) captures nearly all of the available gain: across all 29
languages the 3.0 GB v3 never improves on it by more than
1.2 percentage points. The extra 2 GB buys very little.
3. Bigger is not monotonically better
turbo (1.6 GB) — the model directly above the default in most
pickers, at three times the download — scores numerically worse than the
547 MB default in 9 of 29 languages: Hindi (42.1% vs 34.4%), Thai (24.6% vs 20.8%), Chinese (Simplified) (16.8% vs 13.7%), Malay (9.8% vs 7.1%), Vietnamese (10.3% vs 9.8%), Indonesian (7.8% vs 7.4%), Turkish (15.5% vs 15.2%), Greek (11.3% vs 11.2%), French (10.1% vs 10.0%). (Margins that small are
not all individually significant at this sample size; the point is the direction, which
size alone would not predict.) Reaching for the next model up is not a safe default.
4. tiny is unusable in most of the world
The 75 MB model is frequently reported as a viable lightweight option. Measured across 29 languages it is not: Hindi 100.0%, Danish 87.2%, Hungarian 84.8%, Czech 84.0%, Hebrew 81.9%. It is a placeholder that loads while a real model downloads, not a model to work in.
The hardest and easiest languages
The best result in the entire sweep is Spanish at 2.8%. The worst is Hebrew, which does not fall below 48.3% in any configuration tested — no model choice fixes it. If your work is in that language, local Whisper transcription of any size will need correction.
Related: silent truncation — when a model stops early and the app reports success →
Correction: the Indic and Arabic rows are under re-measurement
Dated 2026-08-28. The scorer behind this table normalised text with
[^\w\s], and Python's \w does not match Unicode combining marks.
Every Devanagari vowel sign, Bengali nukta and Arabic diacritic was therefore replaced by a
space, splitting one word into several. Measured on the same references this page used:
Tamil scored 1,520 units against 437 real words, Hindi 1,237 against 683,
Assamese 1,220 against 498. Latin, Cyrillic and CJK rows are unaffected — those scripts carry
no combining marks — and so are the conclusions that rest on them.
This inflates the reference and, because fragments cannot align, the error count with it. So the rows for Hindi, Arabic, Hebrew and the other combining-mark scripts overstate the error rate by an amount that varies per language, and the “worst language” claim above is not safe for those rows. A corrected sweep across 83 languages is running now and this page will be regenerated from it.
It is recorded here rather than quietly deleted, for the same reason as the retraction below: a normaliser that silently mangles non-Latin script is exactly the failure a reader of a multilingual benchmark should be told about, and it is not visible in any published number.
A retraction, kept in place
An earlier version of this sweep reported large-v3 as significantly better than the
default for Ukrainian. That was wrong. It was measured under beam 5, which is
whisper-cli's default but not what the app ships. Re-run greedy — the
real setting — the default and v3 land in a dead tie. The gap was the
decoder, not the model. It is recorded here rather than quietly deleted, because the
decoder setting is exactly the variable most local-Whisper comparisons leave unstated.
Scope and limits
- One chunk per language (~1,000 words, or ~2,000 characters for the CER languages). Enough to separate large differences; not enough to call small ones. Where this page says a model "holds", read that as not separated at this sample size, not as proven identical.
whisper-cliwith the app's context and decode settings — not the full Lesskeys pipeline, which adds voice-activity detection, a hallucination filter, loudness normalization and a language clamp. Real-world rates in the app are better than these.- Read speech, not meetings. Conversational and overlapping audio is materially harder.
- Absolute rates carry some inflation from number and punctuation formatting differences that a human reader would not count as errors. Comparisons between models on the same language are the trustworthy part; comparisons between languages are noisier.
Free to download. One-time unlock. No subscription. · Pro $49.99 one-time · macOS 13+
FAQ
Which Whisper model should I run on a Mac?
In 23 of the 29 languages measured, the 547 MB quantized Large V3 Turbo model is as accurate as anything larger, so it is the sensible default. The exceptions are listed on this page; in those languages the 1.0 GB v3-q5 model is worth the extra download.
Does quantizing a Whisper model hurt accuracy?
Barely. Across all 29 languages the full-precision 3.0 GB model never beats the 1.0 GB quantized one by more than 1.2 percentage points, and the 547 MB quantized default matches much larger models in most languages.
Is a bigger Whisper model always more accurate?
No. The 1.6 GB turbo model scores numerically worse than the 547 MB default in 9 of the 29 languages measured. Model size is not a reliable proxy for accuracy in a given language.
Why is word error rate not used for Japanese, Korean, Thai and Chinese?
Those languages are not space-delimited, so word-level alignment is not meaningful. Character error rate is the standard substitute and is what this page reports for them.
Was this measured in the cloud?
No. Every number was produced locally on Apple Silicon, which is also how Lesskeys runs transcription — the audio never leaves the machine.