22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.
| Rank | Model | Word Error Rate | Character Error Rate | Word Accuracy | Exact Match Rate | P90 WER | Median Latency |
|---|---|---|---|---|---|---|---|
| 01 | qwen3-asr-flash-2026-02-10Qwen | 1.53%1.33–1.76 | 0.94% | 98.47% | 66.1% | 5.72% | 0.7 s |
| 02 | universal-3-5-proAssemblyAI | 1.55%1.36–1.78 | 0.89% | 98.45% | 65.7% | 5.88% | 0.6 s |
| 03 | mai-transcribe-1.5Microsoft | 1.57%1.37–1.80 | 0.87% | 98.43% | 65.5% | 5.88% | 1.3 s |
| 04 | transcribe-1Fish Audio | 1.59%1.39–1.83 | 0.90% | 98.41% | 64.1% | 6.45% | 0.7 s |
| 05 | qwen3-asr-1.7bQwen | 1.60%1.40–1.84 | 0.91% | 98.40% | 64.5% | 6.45% | 1.8 s |
Accuracy
Error breakdown
Robustness
By clip length
Speed
| |||||||
| 06 | gpt-transcribeOpenAI | 1.76%1.57–1.99 | 1.05% | 98.24% | 61.3% | 5.88% | 1.0 s |
| 07 | mai-transcribe-2Microsoft | 1.76%1.56–1.99 | 1.06% | 98.24% | 61.1% | 6.09% | 0.8 s |
| 08 | grok-stt-1.0xAI | 1.77%1.57–2.01 | 0.99% | 98.23% | 61.1% | 6.97% | 1.2 s |
| 09 | qwen3-asr-0.6bQwen | 1.84%1.61–2.09 | 0.99% | 98.16% | 60.6% | 7.14% | 3.2 s |
| 10 | gpt-4o-mini-transcribeOpenAI | 1.85%1.65–2.08 | 1.09% | 98.15% | 60.5% | 6.67% | 1.0 s |
| 11 | muse-voice-transcribe-1.0Meta | 1.86%1.66–2.11 | 1.02% | 98.14% | 60.2% | 6.67% | 3.7 s |
| 12 | chirp-3Google | 1.90%1.68–2.15 | 1.04% | 98.10% | 59.3% | 6.84% | 5.3 s |
| 13 | nova-3Deepgram | 2.12%1.92–2.37 | 1.21% | 97.88% | 54.0% | 7.14% | 0.4 s |
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 2.13%1.93–2.39 | 1.20% | 97.87% | 54.0% | 6.72% | 0.7 s |
| 15 | voxtral-small-24b-2507-sttMistral AI | 2.14%1.91–2.41 | 1.33% | 97.86% | 57.1% | 7.14% | 2.7 s |
| 16 | voxtral-mini-transcribeMistral AI | 2.15%1.94–2.40 | 1.31% | 97.85% | 55.2% | 7.14% | 1.0 s |
| 17 | whisper-large-v3OpenAI | 2.38%2.14–2.63 | 1.43% | 97.62% | 51.8% | 7.41% | 1.4 s |
| 18 | whisper-1OpenAI | 2.51%2.27–2.77 | 1.55% | 97.49% | 52.4% | 8.00% | 1.8 s |
| 19 | whisper-large-v3-turboOpenAI | 2.55%2.32–2.83 | 1.59% | 97.45% | 49.4% | 8.33% | 3.6 s |
| 20 | gpt-4o-transcribeOpenAI | 3.17%2.86–3.53 | 2.28% | 96.83% | 48.2% | 10.40% | 1.2 s |
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 3.35%3.10–3.63 | 1.88% | 96.65% | 43.6% | 9.75% | 3.8 s |
| 22 | voxtral-mini-3b-2507Mistral AI | 3.89%2.20–6.09 | 3.23% | 96.11% | 55.1% | 7.79% | 1.3 s |
Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Request time and real-time factor. Measured under concurrent load, so indicative only.
We run it on the same audio, under the same scoring, and publish it next to the other 22.