22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.
| Rank | Model | Word Error Rate | WER 95% CI (low) | WER 95% CI (high) | Word Accuracy | Character Error Rate | Strict WER | Match Error Rate | Word Information Lost | Word Information Preserved |
|---|---|---|---|---|---|---|---|---|---|---|
| 22 | voxtral-mini-3b-2507Mistral AI | 3.89%2.20–6.09 | 2.20% | 6.09% | 96.11% | 3.23% | 6.86% | 3.81% | 4.59% | 95.41% |
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 3.35%3.10–3.63 | 3.10% | 3.63% | 96.65% | 1.88% | 6.37% | 3.34% | 4.79% | 95.21% |
| 20 | gpt-4o-transcribeOpenAI | 3.17%2.86–3.53 | 2.86% | 3.53% | 96.83% | 2.28% | 6.21% | 3.16% | 3.85% | 96.15% |
| 19 | whisper-large-v3-turboOpenAI | 2.55%2.32–2.83 | 2.32% | 2.83% | 97.45% | 1.59% | 5.48% | 2.53% | 3.43% | 96.57% |
| 18 | whisper-1OpenAI | 2.51%2.27–2.77 | 2.27% | 2.77% | 97.49% | 1.55% | 5.47% | 2.50% | 3.36% | 96.64% |
| 17 | whisper-large-v3OpenAI | 2.38%2.14–2.63 | 2.14% | 2.63% | 97.62% | 1.43% | 5.34% | 2.37% | 3.25% | 96.75% |
| 16 | voxtral-mini-transcribeMistral AI | 2.15%1.94–2.40 | 1.94% | 2.40% | 97.85% | 1.31% | 5.23% | 2.14% | 2.91% | 97.09% |
| 15 | voxtral-small-24b-2507-sttMistral AI | 2.14%1.91–2.41 | 1.91% | 2.41% | 97.86% | 1.33% | 5.17% | 2.13% | 2.85% | 97.15% |
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 2.13%1.93–2.39 | 1.93% | 2.39% | 97.87% | 1.20% | 4.79% | 2.11% | 2.99% | 97.01% |
| 13 | nova-3Deepgram | 2.12%1.92–2.37 | 1.92% | 2.37% | 97.88% | 1.21% | 5.29% | 2.11% | 2.95% | 97.05% |
| 12 | chirp-3Google | 1.90%1.68–2.15 | 1.68% | 2.15% | 98.10% | 1.04% | 5.00% | 1.89% | 2.72% | 97.28% |
| 11 | muse-voice-transcribe-1.0Meta | 1.86%1.66–2.11 | 1.66% | 2.11% | 98.14% | 1.02% | 4.66% | 1.85% | 2.68% | 97.32% |
| 10 | gpt-4o-mini-transcribeOpenAI | 1.85%1.65–2.08 | 1.65% | 2.08% | 98.15% | 1.09% | 4.91% | 1.84% | 2.55% | 97.45% |
| 09 | qwen3-asr-0.6bQwen | 1.84%1.61–2.09 | 1.61% | 2.09% | 98.16% | 0.99% | 4.78% | 1.83% | 2.63% | 97.37% |
| 08 | grok-stt-1.0xAI | 1.77%1.57–2.01 | 1.57% | 2.01% | 98.23% | 0.99% | 5.00% | 1.76% | 2.54% | 97.46% |
| 07 | mai-transcribe-2Microsoft | 1.76%1.56–1.99 | 1.56% | 1.99% | 98.24% | 1.06% | 4.74% | 1.75% | 2.39% | 97.61% |
| 06 | gpt-transcribeOpenAI | 1.76%1.57–1.99 | 1.57% | 1.99% | 98.24% | 1.05% | 4.84% | 1.75% | 2.42% | 97.58% |
| 05 | qwen3-asr-1.7bQwen | 1.60%1.40–1.84 | 1.40% | 1.84% | 98.40% | 0.91% | 4.52% | 1.59% | 2.25% | 97.75% |
Accuracy
Error breakdown
Robustness
By clip length
Speed
| ||||||||||
| 04 | transcribe-1Fish Audio | 1.59%1.39–1.83 | 1.39% | 1.83% | 98.41% | 0.90% | 4.45% | 1.58% | 2.24% | 97.76% |
| 03 | mai-transcribe-1.5Microsoft | 1.57%1.37–1.80 | 1.37% | 1.80% | 98.43% | 0.87% | 4.24% | 1.56% | 2.18% | 97.82% |
| 02 | universal-3-5-proAssemblyAI | 1.55%1.36–1.78 | 1.36% | 1.78% | 98.45% | 0.89% | 4.46% | 1.54% | 2.16% | 97.84% |
| 01 | qwen3-asr-flash-2026-02-10Qwen | 1.53%1.33–1.76 | 1.33% | 1.76% | 98.47% | 0.94% | 4.50% | 1.53% | 2.04% | 97.96% |
Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Request time and real-time factor. Measured under concurrent load, so indicative only.
We run it on the same audio, under the same scoring, and publish it next to the other 22.