22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.
| Rank | Model | WER - short audio (under 5 s) | WER - medium audio (5 to 15 s) | WER - long audio (over 15 s) |
|---|---|---|---|---|
| 01 | qwen3-asr-flash-2026-02-10Qwen | 3.00% | 1.81% | 1.29% |
| 02 | universal-3-5-proAssemblyAI | 2.76% | 2.01% | 1.22% |
| 03 | mai-transcribe-1.5Microsoft | 2.88% | 1.93% | 1.28% |
| 04 | transcribe-1Fish Audio | 3.70% | 2.05% | 1.22% |
| 05 | qwen3-asr-1.7bQwen | 3.88% | 2.00% | 1.24% |
| 06 | gpt-transcribeOpenAI | 3.29% | 2.19% | 1.43% |
| 07 | mai-transcribe-2Microsoft | 3.41% | 2.06% | 1.50% |
| 08 | grok-stt-1.0xAI | 3.88% | 2.27% | 1.37% |
| 09 | qwen3-asr-0.6bQwen | 4.50% | 2.24% | 1.47% |
| 10 | gpt-4o-mini-transcribeOpenAI | 3.53% | 2.33% | 1.47% |
| 11 | muse-voice-transcribe-1.0Meta | 3.41% | 2.44% | 1.44% |
| 12 | chirp-3Google | 3.64% | 2.49% | 1.46% |
| 13 | nova-3Deepgram | 5.29% | 2.56% | 1.69% |
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 3.76% | 2.30% | 1.95% |
Accuracy
Error breakdown
Robustness
By clip length
Speed
| ||||
| 15 | voxtral-small-24b-2507-sttMistral AI | 3.47% | 2.62% | 1.79% |
| 16 | voxtral-mini-transcribeMistral AI | 4.05% | 2.69% | 1.74% |
| 17 | whisper-large-v3OpenAI | 4.58% | 3.02% | 1.89% |
| 18 | whisper-1OpenAI | 4.47% | 3.00% | 2.12% |
| 19 | whisper-large-v3-turboOpenAI | 4.99% | 3.02% | 2.14% |
| 20 | gpt-4o-transcribeOpenAI | 5.17% | 4.66% | 2.20% |
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 5.76% | 3.79% | 2.97% |
| 22 | voxtral-mini-3b-2507Mistral AI | 33.43% | 4.62% | 1.86% |
Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Request time and real-time factor. Measured under concurrent load, so indicative only.
We run it on the same audio, under the same scoring, and publish it next to the other 22.