Where the bands come from. We did not run these benchmarks ourselves. Each band is a truthful coarsening of a published result: we take the published per-language figure and report only which of the three ranges it falls into. The band is not a new measurement or estimate, and the published result for every language shown falls inside its displayed band. When no result has been published for a language, the row says so.
FLEURS.A public benchmark from Google Research covering 102 languages. Each language's test set is read speech: sentences from the FLoRes translation corpus recorded by native speakers, a few hours of audio per language. Read speech is recorded under cleaner conditions than meetings, lectures, podcasts, and other real-world audio, so treat the bands as a best case.
Scoring. Figures behind the bands are word error rates, except for languages without whitespace word boundaries (Chinese, Japanese, Cantonese), which are scored per character in the cited evaluation, as is standard practice. WER also depends heavily on how text is normalized before scoring: lowercasing, punctuation removal, number and abbreviation expansion. Normalization differences alone can move WER by whole percentage points, which is one reason accuracy claims without stated normalization rules are not comparable, and why we do not compare results across languages.
Comparisons. We only compare models when the figures come from the same published evaluation on the same splits. Where such comparisons exist, the engine we ship shows 1.5x-2.3x fewer errors than Whisper large-v3 in published same-split comparisons. We do not publish per-language ratios or estimate results a vendor has not published.
What transcribe.so ships. Every file runs on our Standard engine, the production stack summarized above, with speaker labels and word-level timestamps on every transcript. Diarization is included in the base price, not an add-on.