Ranked by word error rate. Switch tabs to see accuracy, the kind of error, robustness on hard clips, clip length, and what each model costs to run.
| Rank | Model | Word Error Rate | WER 95% CI (low) | WER 95% CI (high) | Word Accuracy | Character Error Rate | Strict WER | Match Error Rate | Word Information Lost | Word Information Preserved |
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | qwen3-asr-flash-2026-02-10Qwen | 0.91%0.56–1.33 | 0.56% | 1.33% | 99.09% | 0.50% | 4.60% | 0.91% | 1.22% | 98.78% |
| 02 | transcribe-1Fish Audio | 1.28%0.79–1.83 | 0.79% | 1.83% | 98.72% | 0.74% | 4.97% | 1.27% | 1.78% | 98.22% |
| 03 | mai-transcribe-1.5Microsoft | 1.33%0.88–1.84 | 0.88% | 1.84% | 98.67% | 0.63% | 4.89% | 1.33% | 1.89% | 98.11% |
| 04 | qwen3-asr-0.6bQwen | 1.36%0.90–1.94 | 0.90% | 1.94% | 98.64% | 0.64% | 5.03% | 1.36% | 2.12% | 97.88% |
| 05 | qwen3-asr-1.7bQwen | 1.36%0.87–1.86 | 0.87% | 1.86% | 98.64% | 0.75% | 5.06% | 1.36% | 1.92% | 98.08% |
| 06 | chirp-3Google | 1.45%1.01–1.94 | 1.01% | 1.94% | 98.55% | 0.78% | 5.40% | 1.44% | 2.18% | 97.82% |
| 07 | mai-transcribe-2Microsoft | 1.45%0.99–1.94 | 0.99% | 1.94% | 98.55% | 0.80% | 5.14% | 1.44% | 2.01% | 97.99% |
| 08 | universal-3-5-proAssemblyAI | 1.48%1.01–1.94 | 1.01% | 1.94% | 98.52% | 0.70% | 5.03% | 1.47% | 2.12% | 97.88% |
| 09 | muse-voice-transcribe-1.0Meta | 1.48%1.00–2.11 | 1.00% | 2.11% | 98.52% | 0.75% | 5.03% | 1.47% | 2.09% | 97.91% |
| 10 | gpt-transcribeOpenAI | 1.59%1.07–2.15 | 1.07% | 2.15% | 98.41% | 0.91% | 5.37% | 1.59% | 2.09% | 97.91% |
| 11 | grok-stt-1.0xAI | 1.65%1.19–2.20 | 1.19% | 2.20% | 98.35% | 0.83% | 5.26% | 1.64% | 2.51% | 97.49% |
| 12 | gpt-4o-mini-transcribeOpenAI | 1.82%1.22–2.42 | 1.22% | 2.42% | 98.18% | 1.06% | 5.46% | 1.81% | 2.52% | 97.48% |
| 13 | voxtral-mini-transcribeMistral AI | 1.96%1.42–2.52 | 1.42% | 2.52% | 98.04% | 1.07% | 5.86% | 1.95% | 2.66% | 97.34% |
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 1.99%1.45–2.55 | 1.45% | 2.55% | 98.01% | 1.03% | 5.46% | 1.97% | 2.84% | 97.16% |
| 15 | nova-3Deepgram | 2.06%1.55–2.64 | 1.55% | 2.64% | 97.94% | 1.19% | 5.83% | 2.06% | 2.84% | 97.16% |
| 16 | whisper-1OpenAI | 2.13%1.55–2.76 | 1.55% | 2.76% | 97.87% | 1.22% | 5.72% | 2.12% | 2.85% | 97.15% |
| 17 | voxtral-mini-3b-2507Mistral AI | 2.24%1.56–3.04 | 1.56% | 3.04% | 97.76% | 1.43% | 6.09% | 2.23% | 2.97% | 97.03% |
| 18 | voxtral-small-24b-2507-sttMistral AI | 2.27%1.60–3.08 | 1.60% | 3.08% | 97.73% | 1.32% | 6.09% | 2.27% | 3.11% | 96.89% |
| 19 | whisper-large-v3-turboOpenAI | 2.33%1.76–3.04 | 1.76% | 3.04% | 97.67% | 1.42% | 5.89% | 2.32% | 3.30% | 96.70% |
| 20 | whisper-large-v3OpenAI | 2.35%1.75–3.08 | 1.75% | 3.08% | 97.65% | 1.49% | 6.26% | 2.35% | 3.19% | 96.81% |
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 3.12%2.31–4.08 | 2.31% | 4.08% | 96.88% | 1.59% | 6.83% | 3.11% | 4.48% | 95.52% |
| 22 | gpt-4o-transcribeOpenAI | 3.94%2.64–5.75 | 2.64% | 5.75% | 96.06% | 2.92% | 7.63% | 3.93% | 4.64% | 95.36% |
Rates are pooled across all clips unless the name says mean, median or P90. Every metric is computed after the same normalization on both the reference and the model output.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Billed cost and request time. Latency was measured under concurrent load and is indicative only.
We run it on the same audio, through the same normalization, and publish it next to the other 22.