22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.
| Rank | Model | Median Latency | P95 Latency | Real-Time Factor (median) |
|---|---|---|---|---|
| 01 | qwen3-asr-flash-2026-02-10Qwen | 0.7 s | 1.3 s | 0.06× |
| 02 | universal-3-5-proAssemblyAI | 0.6 s | 1.3 s | 0.06× |
| 03 | mai-transcribe-1.5Microsoft | 1.3 s | 1.7 s | 0.12× |
| 04 | transcribe-1Fish Audio | 0.7 s | 1.5 s | 0.06× |
| 05 | qwen3-asr-1.7bQwen | 1.8 s | 12.2 s | 0.15× |
| 06 | gpt-transcribeOpenAI | 1.0 s | 1.8 s | 0.09× |
| 07 | mai-transcribe-2Microsoft | 0.8 s | 1.1 s | 0.07× |
| 08 | grok-stt-1.0xAI | 1.2 s | 3.4 s | 0.10× |
| 09 | qwen3-asr-0.6bQwen | 3.2 s | 16.8 s | 0.27× |
| 10 | gpt-4o-mini-transcribeOpenAI | 1.0 s | 1.8 s | 0.09× |
| 11 | muse-voice-transcribe-1.0Meta | 3.7 s | 6.9 s | 0.30× |
| 12 | chirp-3Google | 5.3 s | 11.9 s | 0.46× |
| 13 | nova-3Deepgram | 0.4 s | 1.3 s | 0.04× |
| 14 | parakeet-tdt-0.6b-v3NVIDIA | 0.7 s | 2.1 s | 0.06× |
| 15 | voxtral-small-24b-2507-sttMistral AI | 2.7 s | 5.2 s | 0.23× |
| 16 | voxtral-mini-transcribeMistral AI | 1.0 s | 2.0 s | 0.09× |
| 17 | whisper-large-v3OpenAI | 1.4 s | 3.8 s | 0.12× |
| 18 | whisper-1OpenAI | 1.8 s | 3.3 s | 0.14× |
| 19 | whisper-large-v3-turboOpenAI | 3.6 s | 16.7 s | 0.29× |
| 20 | gpt-4o-transcribeOpenAI | 1.2 s | 2.0 s | 0.10× |
Accuracy
Error breakdown
Robustness
By clip length
Speed
| ||||
| 21 | nemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA | 3.8 s | 7.1 s | 0.32× |
| 22 | voxtral-mini-3b-2507Mistral AI | 1.3 s | 2.8 s | 0.11× |
Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Request time and real-time factor. Measured under concurrent load, so indicative only.
We run it on the same audio, under the same scoring, and publish it next to the other 22.