Ranked by word error rate. Switch tabs to see accuracy, the kind of error, robustness on hard clips, clip length, and what each model costs to run.
| Rank | Model | Word Error Rate | Character Error Rate | Word Accuracy | Exact Match Rate | Price per 1,000 min | Median Latency |
|---|---|---|---|---|---|---|---|
| 01 | qwen3-asr-flash-2026-02-10Qwen | 0.91%0.56–1.33 | 0.50% | 99.09% | 78.0% | $2.03 | 9.3 s |
| 02 | transcribe-1Fish Audio | 1.28%0.79–1.83 | 0.74% | 98.72% | 70.0% | $6.27 | 7.1 s |
| 03 | mai-transcribe-1.5Microsoft | 1.33%0.88–1.84 | 0.63% | 98.67% | 71.0% | $6.27 | 7.5 s |
| 04 | qwen3-asr-0.6bQwen | 1.36%0.90–1.94 | 0.64% | 98.64% | 71.0% | $0.20 | 6.7 s |
| 05 | qwen3-asr-1.7bQwen | 1.36%0.87–1.86 | 0.75% | 98.64% | 70.0% | $0.45 | 5.8 s |
| 06 | chirp-3Google | 1.45%1.01–1.94 | 0.78% | 98.55% | 65.0% | $16.72 | 9.9 s |
| 07 | mai-transcribe-2Microsoft | 1.45%0.99–1.94 | 0.80% | 98.55% | 69.0% | $1.74 | 7.8 s |
| 08 | universal-3-5-proAssemblyAI | 1.48%1.01–1.94 | 0.70% | 98.52% | 67.0% | $3.75 | 10.3 s |
| 09 | muse-voice-transcribe-1.0Meta | 1.48%1.00–2.11 | 0.75% | 98.52% | 66.0% | $3.01 | 10.0 s |
| 10 | gpt-transcribeOpenAI | 1.59%1.07–2.15 | 0.91% | 98.41% | 68.0% | $4.70 | 8.3 s |
| 11 | grok-stt-1.0xAI | 1.65%1.19–2.20 | 0.83% | 98.35% | 62.0% | $1.67 | 7.5 s |
| 12 | gpt-4o-mini-transcribeOpenAI | 1.82%1.22–2.42 | 1.06% | 98.18% | 66.0% | $1.75 | 10.5 s |
| 13 | voxtral-mini-transcribeMistral AI | 1.96%1.42–2.52 | 1.07% | 98.04% | 55.0% | $2.89 | 8.6 s |
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 1.99%1.45–2.55 | 1.03% | 98.01% | 61.0% | $1.50 | 7.0 s |
| 15 | nova-3Deepgram | 2.06%1.55–2.64 | 1.19% | 97.94% | 51.5% | $4.30 | 8.0 s |
| 16 | whisper-1OpenAI | 2.13%1.55–2.76 | 1.22% | 97.87% | 58.0% | $6.27 | 8.3 s |
| 17 | voxtral-mini-3b-2507Mistral AI | 2.24%1.56–3.04 | 1.43% | 97.76% | 57.0% | $1.00 | 7.3 s |
| 18 | voxtral-small-24b-2507-sttMistral AI | 2.27%1.60–3.08 | 1.32% | 97.73% | 53.0% | $3.00 | 9.4 s |
| 19 | whisper-large-v3-turboOpenAI | 2.33%1.76–3.04 | 1.42% | 97.67% | 49.0% | $0.23 | 8.5 s |
| 20 | whisper-large-v3OpenAI | 2.35%1.75–3.08 | 1.49% | 97.65% | 50.0% | $1.59 | 7.6 s |
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 3.12%2.31–4.08 | 1.59% | 96.88% | 52.0% | $0.20 | 10.0 s |
Accuracy
Error breakdown
Robustness
By clip length
Cost & speed
| |||||||
| 22 | gpt-4o-transcribeOpenAI | 3.94%2.64–5.75 | 2.92% | 96.06% | 47.0% | $3.44 | 6.8 s |
Rates are pooled across all clips unless the name says mean, median or P90. Every metric is computed after the same normalization on both the reference and the model output.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Billed cost and request time. Latency was measured under concurrent load and is indicative only.
We run it on the same audio, through the same normalization, and publish it next to the other 22.