22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.
| Rank | Model | Substitution Rate | Deletion Rate | Insertion Rate | Sentence Error Rate | Exact Match Rate | Length Ratio |
|---|---|---|---|---|---|---|---|
| 01 | qwen3-asr-flash-2026-02-10Qwen | 0.51% | 0.55% | 0.47% | 33.9% | 66.1% | 0.999× |
| 02 | universal-3-5-proAssemblyAI | 0.62% | 0.40% | 0.54% | 34.3% | 65.7% | 1.001× |
| 03 | mai-transcribe-1.5Microsoft | 0.63% | 0.26% | 0.68% | 34.5% | 65.5% | 1.004× |
| 04 | transcribe-1Fish Audio | 0.66% | 0.36% | 0.58% | 35.9% | 64.1% | 1.002× |
| 05 | qwen3-asr-1.7bQwen | 0.67% | 0.35% | 0.58% | 35.5% | 64.5% | 1.002× |
| 06 | gpt-transcribeOpenAI | 0.67% | 0.66% | 0.43% | 38.7% | 61.3% | 0.998× |
| 07 | mai-transcribe-2Microsoft | 0.65% | 0.51% | 0.60% | 38.9% | 61.1% | 1.001× |
| 08 | grok-stt-1.0xAI | 0.78% | 0.35% | 0.64% | 38.9% | 61.1% | 1.003× |
| 09 | qwen3-asr-0.6bQwen | 0.81% | 0.44% | 0.59% | 39.4% | 60.6% | 1.001× |
| 10 | gpt-4o-mini-transcribeOpenAI | 0.72% | 0.56% | 0.56% | 39.5% | 60.5% | 1.000× |
| 11 | muse-voice-transcribe-1.0Meta | 0.84% | 0.46% | 0.56% | 39.8% | 60.2% | 1.001× |
| 12 | chirp-3Google | 0.85% | 0.39% | 0.66% | 40.8% | 59.3% | 1.003× |
Accuracy
Error breakdown
Robustness
By clip length
Speed
| |||||||
| 13 | nova-3Deepgram | 0.85% | 0.72% | 0.55% | 46.0% | 54.0% | 0.998× |
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 0.89% | 0.30% | 0.94% | 46.0% | 54.0% | 1.006× |
| 15 | voxtral-small-24b-2507-sttMistral AI | 0.73% | 0.92% | 0.49% | 42.9% | 57.1% | 0.996× |
| 16 | voxtral-mini-transcribeMistral AI | 0.77% | 0.92% | 0.47% | 44.8% | 55.2% | 0.996× |
| 17 | whisper-large-v3OpenAI | 0.89% | 0.95% | 0.54% | 48.2% | 51.8% | 0.996× |
| 18 | whisper-1OpenAI | 0.87% | 1.07% | 0.57% | 47.6% | 52.4% | 0.995× |
| 19 | whisper-large-v3-turboOpenAI | 0.90% | 0.92% | 0.72% | 50.6% | 49.4% | 0.998× |
| 20 | gpt-4o-transcribeOpenAI | 0.69% | 2.12% | 0.36% | 51.8% | 48.2% | 0.982× |
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 1.48% | 1.43% | 0.44% | 56.4% | 43.6% | 0.990× |
| 22 | voxtral-mini-3b-2507Mistral AI | 0.78% | 0.96% | 2.15% | 44.9% | 55.1% | 1.012× |
Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Request time and real-time factor. Measured under concurrent load, so indicative only.
We run it on the same audio, under the same scoring, and publish it next to the other 22.