22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.
| Rank | Model | Mean WER | Median WER | P90 WER | Severe Error Rate | Empty Output Rate | Success Rate |
|---|---|---|---|---|---|---|---|
| 22 | voxtral-mini-3b-2507Mistral AI | 10.64% | 0.00% | 7.79% | 0.5% | 0.0% | 99.8% |
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 8.57% | 2.22% | 9.75% | 0.6% | 0.3% | 99.9% |
| 20 | gpt-4o-transcribeOpenAI | 8.59% | 1.28% | 10.40% | 1.0% | 0.0% | 100.0% |
| 19 | whisper-large-v3-turboOpenAI | 7.87% | 1.23% | 8.33% | 0.7% | 0.0% | 100.0% |
| 18 | whisper-1OpenAI | 7.66% | 0.00% | 8.00% | 0.5% | 0.0% | 100.0% |
| 17 | whisper-large-v3OpenAI | 7.41% | 0.00% | 7.41% | 0.2% | 0.1% | 100.0% |
| 16 | voxtral-mini-transcribeMistral AI | 7.07% | 0.00% | 7.14% | 0.1% | 0.0% | 100.0% |
| 15 | voxtral-small-24b-2507-sttMistral AI | 7.09% | 0.00% | 7.14% | 0.3% | 0.0% | 99.9% |
Accuracy
Error breakdown
Robustness
By clip length
Speed
| |||||||
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 7.05% | 0.00% | 6.72% | 0.3% | 0.1% | 100.0% |
| 13 | nova-3Deepgram | 7.20% | 0.00% | 7.14% | 0.3% | 0.0% | 99.8% |
| 12 | chirp-3Google | 7.26% | 0.00% | 6.84% | 0.6% | 0.1% | 100.0% |
| 11 | muse-voice-transcribe-1.0Meta | 6.90% | 0.00% | 6.67% | 0.4% | 0.1% | 100.0% |
| 10 | gpt-4o-mini-transcribeOpenAI | 6.98% | 0.00% | 6.67% | 0.3% | 0.0% | 100.0% |
| 09 | qwen3-asr-0.6bQwen | 8.28% | 0.00% | 7.14% | 1.1% | 0.0% | 91.0% |
| 08 | grok-stt-1.0xAI | 6.81% | 0.00% | 6.97% | 0.2% | 0.0% | 100.0% |
| 07 | mai-transcribe-2Microsoft | 6.62% | 0.00% | 6.09% | 0.2% | 0.0% | 100.0% |
| 06 | gpt-transcribeOpenAI | 6.64% | 0.00% | 5.88% | 0.3% | 0.0% | 100.0% |
| 05 | qwen3-asr-1.7bQwen | 7.41% | 0.00% | 6.45% | 1.1% | 0.0% | 100.0% |
| 04 | transcribe-1Fish Audio | 7.40% | 0.00% | 6.45% | 1.0% | 0.0% | 100.0% |
| 03 | mai-transcribe-1.5Microsoft | 6.38% | 0.00% | 5.88% | 0.1% | 0.0% | 100.0% |
| 02 | universal-3-5-proAssemblyAI | 6.38% | 0.00% | 5.88% | 0.2% | 0.0% | 100.0% |
| 01 | qwen3-asr-flash-2026-02-10Qwen | 6.86% | 0.00% | 5.72% | 0.7% | 0.0% | 100.0% |
Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Request time and real-time factor. Measured under concurrent load, so indicative only.
We run it on the same audio, under the same scoring, and publish it next to the other 22.