Accuracy benchmarks

Transcription accuracy by language

Published error-rate bands for the engine transcribe.so ships, per language. Each band is a truthful coarsening of a published benchmark result: the published figure for every language shown falls inside its displayed band. Nothing on this page is estimated, and we say "Not published" rather than guess.

Last updated 2026-07-28. Jump to: Korean · Japanese · Chinese (Mandarin) · English · Methodology · Limitations

How to read this page

WER (word error rate) is the standard accuracy metric for speech-to-text. It counts wrong, missing, and extra words against a human-verified reference transcript, so lower is better. For languages written without spaces, such as Chinese, Japanese, and Cantonese, the same count is done per character in the cited evaluation. We group each published result into one of three bands: Under 5%, 5% to under 15%, and 15% or higher. The same boundaries apply to character error rate for the languages scored per character. Bands are comparable within a language, not across languages.

Error-rate band by language

FLEURS is a public multilingual benchmark of read speech built by Google Research. The table lists every language our Standard engine, the default engine on transcribe.so, supports, in alphabetical order, with the band its published FLEURS result falls into. The bands summarize published results; they are not our own measurements.

LanguagePublished error-rate bandFLEURS, lower is better
Arabicالعربية5% to under 15%
Cantonese粵語Under 5%
Chinese (Mandarin)中文Under 5%
CzechČeština15% or higher
DanishDansk5% to under 15%
DutchNederlandsUnder 5%
EnglishEnglishUnder 5%
FilipinoFilipino15% or higher
FinnishSuomi5% to under 15%
FrenchFrançaisUnder 5%
GermanDeutschUnder 5%
GreekΕλληνικά5% to under 15%
Hindiहिन्दी5% to under 15%
HungarianMagyar15% or higher
IndonesianBahasa IndonesiaUnder 5%
ItalianItalianoUnder 5%
Japanese日本語Under 5%
Korean한국어Under 5%
MacedonianМакедонскиNot published
MalayBahasa Melayu5% to under 15%
Persianفارسی15% or higher
PolishPolski5% to under 15%
PortuguesePortuguêsUnder 5%
RomanianRomână5% to under 15%
RussianРусскийUnder 5%
SpanishEspañolUnder 5%
SwedishSvenska15% or higher
Thaiภาษาไทย5% to under 15%
TurkishTürkçe5% to under 15%
VietnameseTiếng ViệtUnder 5%

Bands coarsen results from the engine's published technical report on the FLEURS benchmark (Google Research). Languages are sorted alphabetically, never by result, and we do not publish ordering within a band. Where a same-split comparison has been published, the engine we ship shows 1.5x-2.3x fewer errors than Whisper large-v3 in published same-split comparisons.

See it in action

Bands are one thing. See the actual output.

Error rates tell you how many words land. This is what a finished transcript looks like: chapters, speaker-aware sections, and searchable Q&A on a real recording.

how to find work you *actually* enjoy
Ali Abdaal
Try this real transcript
Who said what, when
Contents✍️Human written like chapters
18 chapters · 59 sections
1Why you're secretly miserable at your 'dream job'
2The shame of hating your paycheck
3Are you built for the pathless path?
4Build confidence before you quit
5The math that makes quitting less scary
Ask this video
🎯Get who said what, when
Answer
Paul left because the work had quietly stopped fitting who he was, not because of a single dramatic event. Early on he chased prestige and big salaries, optimizing for impressive internships and the markers of success . By around thirty-two the job had drained his energy and passion, and quitting was mostly about escaping that misalignment and getting himself back . When he ran a self-assessment, he realized he'd drifted from the goals he set in grad school, to avoid becoming money-obsessed and to keep his sense of humor, which made clear how far off course he'd gone . The decision was less “follow your dream” and more “stop betraying your own values.”

Command Palette

Search for a command to run...

Language details

Korean한국어

Korean sits in the strongest published band, under 5% error rate on FLEURS. That is the territory where a transcript of clean speech is usable without a correction pass.

OpenAI has not published per-language FLEURS results for Whisper large-v3, so we show no Whisper band for Korean rather than estimating one. Korean services built on Whisper-family models inherit whatever Whisper does on Korean, and none of the Korean transcription products we are aware of publish a measured error rate at all.

On transcribe.so, Korean audio runs through the Standard engine by default. Word-level timestamps are supported, and diarization is included in the base price.

Japanese日本語

Japanese also falls in the under 5% band on FLEURS with the Standard engine, which every Japanese file on transcribe.so runs through by default.

Japanese error rates are computed at character level because Japanese text has no whitespace word boundaries. Character-level and word-level percentages are not directly comparable across languages; compare models within a language instead.

As with Korean, OpenAI publishes no per-language FLEURS figure for Whisper large-v3 on Japanese, so no Whisper comparison is shown. We do not estimate results a vendor has not published.

Chinese (Mandarin)中文

Chinese is one of the few languages with a clean same-split Whisper comparison. Across the languages where such comparisons exist, the engine we ship shows 1.5x-2.3x fewer errors than Whisper large-v3 in published same-split comparisons. Mandarin sits in the under 5% band.

Cantonese is also in the under 5% band for our Standard engine, and it is one of the languages covered by the same published same-split comparison. Whisper-derived tools inherit a substantially higher published error rate there.

Chinese dialect and accent coverage is included as part of the 52 languages and dialects transcribe.so supports.

EnglishEnglish

English has a clean same-split comparison: on FLEURS read speech, the engine we ship shows 1.5x-2.3x fewer errors than Whisper large-v3 in published same-split comparisons, and English sits in the under 5% band.

FLEURS is read speech recorded in quiet conditions. Expect absolute error rates on real-world recordings such as meetings, earnings calls, and noisy web audio to be higher for every model; the relative ordering between models is the informative part.

Methodology

Where the bands come from. We did not run these benchmarks ourselves. Each band is a truthful coarsening of a published result: we take the published per-language figure and report only which of the three ranges it falls into. The band is not a new measurement or estimate, and the published result for every language shown falls inside its displayed band. When no result has been published for a language, the row says so.

FLEURS.A public benchmark from Google Research covering 102 languages. Each language's test set is read speech: sentences from the FLoRes translation corpus recorded by native speakers, a few hours of audio per language. Read speech is recorded under cleaner conditions than meetings, lectures, podcasts, and other real-world audio, so treat the bands as a best case.

Scoring. Figures behind the bands are word error rates, except for languages without whitespace word boundaries (Chinese, Japanese, Cantonese), which are scored per character in the cited evaluation, as is standard practice. WER also depends heavily on how text is normalized before scoring: lowercasing, punctuation removal, number and abbreviation expansion. Normalization differences alone can move WER by whole percentage points, which is one reason accuracy claims without stated normalization rules are not comparable, and why we do not compare results across languages.

Comparisons. We only compare models when the figures come from the same published evaluation on the same splits. Where such comparisons exist, the engine we ship shows 1.5x-2.3x fewer errors than Whisper large-v3 in published same-split comparisons. We do not publish per-language ratios or estimate results a vendor has not published.

What transcribe.so ships. Every file runs on our Standard engine, the production stack summarized above, with speaker labels and word-level timestamps on every transcript. Diarization is included in the base price, not an add-on.

Honest limitations

  • FLEURS is read speech recorded in quiet conditions. Real meetings, lectures, and podcasts are noisier and more spontaneous, so absolute error rates on your audio will be higher for every model, and a band on this page does not predict the error rate on your recording.
  • The table leans on one corpus. A model can land differently on other test sets, accents, or domains.
  • Bands deliberately discard precision. Two languages in the same band can have meaningfully different published results; we accept that loss rather than publish exact third-party figures as if they were our own.
  • Whisper large-v3 has no vendor-published per-language FLEURS results for most languages, including Korean and Japanese. We show no comparison there instead of estimating one.
  • Benchmarks do not cover heavily accented speech, domain jargon (medical, legal), overlapping speakers, or very low-quality recordings. No public benchmark we are aware of does this well across languages.
  • Accuracy claims you may see elsewhere, such as "99% accurate", are usually published without a test set, normalization rules, or model version, and cannot be compared to anything on this page.

Test it on your own audio

The only benchmark that matters is your recording. Free credits included, no credit card required, and you see the exact price before you confirm.

Works on YouTube links, podcasts, meetings, lectures, and voice memos.