PERIT-STT · Leaderboard

Speech recognition, measured against human experts

22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.

RankModelSubstitution RateDeletion RateInsertion RateSentence Error RateExact Match RateLength Ratio
22
0.78%
0.96%
2.15%
44.9%
55.1%
1.012×
21
1.48%
1.43%
0.44%
56.4%
43.6%
0.990×
20
0.69%
2.12%
0.36%
51.8%
48.2%
0.982×
19
0.90%
0.92%
0.72%
50.6%
49.4%
0.998×
18
0.87%
1.07%
0.57%
47.6%
52.4%
0.995×
17
0.89%
0.95%
0.54%
48.2%
51.8%
0.996×
16
0.77%
0.92%
0.47%
44.8%
55.2%
0.996×
15voxtral-small-24b-2507-sttMistral AI
0.73%
0.92%
0.49%
42.9%
57.1%
0.996×
Accuracy
Word Error Rate (WER)
2.14%#15/22
WER 95% CI (low)
1.91%#13/22
WER 95% CI (high)
2.41%#16/22
Word Accuracy
97.86%#15/22
Character Error Rate (CER)
1.33%#16/22
Strict WER
5.17%#14/22
Match Error Rate (MER)
2.13%#15/22
Word Information Lost (WIL)
2.85%#13/22
Word Information Preserved (WIP)
97.15%#13/22
Error breakdown
Substitution Rate
0.73%#10/22
Deletion Rate
0.92%#15/22
Insertion Rate
0.49%#6/22
Sentence Error Rate (SER)
42.9%#13/22
Exact Match Rate
57.1%#13/22
Length Ratio
0.996×#14/22
Robustness
Mean WER
7.09%#11/22
Median WER
0.00%#1/22
P90 WER
7.14%#13/22
Severe Error Rate
0.3%#7/22
Empty Output Rate
0.0%#1/22
Success Rate
99.9%#18/22
By clip length
WER - short audio (under 5 s)
3.47%#7/22
WER - medium audio (5 to 15 s)
2.62%#15/22
WER - long audio (over 15 s)
1.79%#15/22
Speed
Median Latency
2.7 s#17/22
P95 Latency
5.2 s#16/22
Real-Time Factor (median)
0.23×#17/22
14
0.89%
0.30%
0.94%
46.0%
54.0%
1.006×
13
0.85%
0.72%
0.55%
46.0%
54.0%
0.998×
12
0.85%
0.39%
0.66%
40.8%
59.3%
1.003×
11
0.84%
0.46%
0.56%
39.8%
60.2%
1.001×
10
0.72%
0.56%
0.56%
39.5%
60.5%
1.000×
09
0.81%
0.44%
0.59%
39.4%
60.6%
1.001×
08
0.78%
0.35%
0.64%
38.9%
61.1%
1.003×
07
0.65%
0.51%
0.60%
38.9%
61.1%
1.001×
06
0.67%
0.66%
0.43%
38.7%
61.3%
0.998×
05
0.67%
0.35%
0.58%
35.5%
64.5%
1.002×
04
0.66%
0.36%
0.58%
35.9%
64.1%
1.002×
03
0.63%
0.26%
0.68%
34.5%
65.5%
1.004×
02
0.62%
0.40%
0.54%
34.3%
65.7%
1.001×
01
0.51%
0.55%
0.47%
33.9%
66.1%
0.999×
Blue marks the best value in each column; a dash means the metric does not apply to that model. Open a model for all 27 metrics and where it places on each.Showing 22 of 22
Best on each axis

The winners, one question at a time.

Most accurate
1.53%Word Error Rate (WER)qwen3-asr-flash-2026-02-10Qwen
Lowest character error
0.87%Character Error Rate (CER)mai-transcribe-1.5Microsoft
Most exact matches
66.1%Exact Match Rateqwen3-asr-flash-2026-02-10Qwen
Most robust
5.72%P90 WERqwen3-asr-flash-2026-02-10Qwen
Fewest dropped words
0.26%Deletion Ratemai-transcribe-1.5Microsoft
Fewest invented words
0.36%Insertion Rategpt-4o-transcribeOpenAI
Fastest
0.4 sMedian Latencynova-3Deepgram
Glossary

What every column means.

Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.

Accuracy

Word- and character-level agreement with the reference, after normalization — and without it.

Word Error Rate (WER)% · lower is better
Headline metric. (substituted + deleted + inserted words) / words in the ground truth, after normalization. Pooled across all clips.
WER 95% CI (low)% · lower is better
Lower bound of the 95% bootstrap confidence interval for WER.
WER 95% CI (high)% · lower is better
Upper bound of the 95% bootstrap confidence interval for WER.
Word Accuracy% · higher is better
100 − WER (floored at 0).
Character Error Rate (CER)% · lower is better
Same as WER but counted on characters; more forgiving of near-miss spellings.
Strict WER% · lower is better
WER with minimal normalization (only lowercase and punctuation removal). Shows how closely raw output matches the ground-truth style.
Match Error Rate (MER)% · lower is better
Errors / (errors + correct words). Bounded at 100% even with heavy hallucination.
Word Information Lost (WIL)% · lower is better
Share of word-level information lost between ground truth and output.
Word Information Preserved (WIP)% · higher is better
100 − WIL.
Error breakdown

What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.

Substitution Rate% · lower is better
Wrong words, as a share of ground-truth words.
Deletion Rate% · lower is better
Missed words (omissions), as a share of ground-truth words.
Insertion Rate% · lower is better
Extra words not spoken (fabrications / hallucinations), as a share of ground-truth words.
Sentence Error Rate (SER)% · lower is better
Share of clips with at least one word error.
Exact Match Rate% · higher is better
Share of clips transcribed perfectly after normalization.
Length Ratiox · closer to 1 is better
Output word count / ground-truth word count. Above 1 suggests hallucination, below 1 suggests skipped speech.
Robustness

How the error is spread across clips — the typical one, the hardest tenth, the outright failures.

Mean WER% · lower is better
Average of per-clip WER (every clip weighted equally).
Median WER% · lower is better
WER of the typical clip.
P90 WER% · lower is better
WER on the hardest 10% of clips; measures robustness.
Severe Error Rate% · lower is better
Share of clips with WER of 50% or more.
Empty Output Rate% · lower is better
Share of clips where the model returned no words.
Success Rate% · higher is better
Share of requests that returned a transcript.
By clip length

WER split by clip duration. Short clips give a model the least context to recover from.

WER - short audio (under 5 s)% · lower is better
WER on short clips.
WER - medium audio (5 to 15 s)% · lower is better
WER on medium-length clips.
WER - long audio (over 15 s)% · lower is better
WER on long clips.
Speed

Request time and real-time factor. Measured under concurrent load, so indicative only.

Median Latencys · lower is better
Median request time. Measured under concurrent load; indicative only.
P95 Latencys · lower is better
95th percentile request time. Measured under concurrent load; indicative only.
Real-Time Factor (median)x · lower is better
Processing time / audio duration. Measured under concurrent load; indicative only.

Want your model on this board?

We run it on the same audio, under the same scoring, and publish it next to the other 22.