PERIT-STT · Leaderboard

Speech recognition, measured against human experts

22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.

RankModelWord Error RateWER 95% CI (low)WER 95% CI (high)Word AccuracyCharacter Error RateStrict WERMatch Error RateWord Information LostWord Information Preserved
01
1.53%1.331.76
1.33%
1.76%
98.47%
0.94%
4.50%
1.53%
2.04%
97.96%
02
1.55%1.361.78
1.36%
1.78%
98.45%
0.89%
4.46%
1.54%
2.16%
97.84%
03
1.57%1.371.80
1.37%
1.80%
98.43%
0.87%
4.24%
1.56%
2.18%
97.82%
04
1.59%1.391.83
1.39%
1.83%
98.41%
0.90%
4.45%
1.58%
2.24%
97.76%
05
1.60%1.401.84
1.40%
1.84%
98.40%
0.91%
4.52%
1.59%
2.25%
97.75%
06
1.76%1.571.99
1.57%
1.99%
98.24%
1.05%
4.84%
1.75%
2.42%
97.58%
07
1.76%1.561.99
1.56%
1.99%
98.24%
1.06%
4.74%
1.75%
2.39%
97.61%
08
1.77%1.572.01
1.57%
2.01%
98.23%
0.99%
5.00%
1.76%
2.54%
97.46%
09
1.84%1.612.09
1.61%
2.09%
98.16%
0.99%
4.78%
1.83%
2.63%
97.37%
10
1.85%1.652.08
1.65%
2.08%
98.15%
1.09%
4.91%
1.84%
2.55%
97.45%
11
1.86%1.662.11
1.66%
2.11%
98.14%
1.02%
4.66%
1.85%
2.68%
97.32%
12
1.90%1.682.15
1.68%
2.15%
98.10%
1.04%
5.00%
1.89%
2.72%
97.28%
13
2.12%1.922.37
1.92%
2.37%
97.88%
1.21%
5.29%
2.11%
2.95%
97.05%
14
2.13%1.932.39
1.93%
2.39%
97.87%
1.20%
4.79%
2.11%
2.99%
97.01%
15
2.14%1.912.41
1.91%
2.41%
97.86%
1.33%
5.17%
2.13%
2.85%
97.15%
16
2.15%1.942.40
1.94%
2.40%
97.85%
1.31%
5.23%
2.14%
2.91%
97.09%
17
2.38%2.142.63
2.14%
2.63%
97.62%
1.43%
5.34%
2.37%
3.25%
96.75%
18
2.51%2.272.77
2.27%
2.77%
97.49%
1.55%
5.47%
2.50%
3.36%
96.64%
19
2.55%2.322.83
2.32%
2.83%
97.45%
1.59%
5.48%
2.53%
3.43%
96.57%
20gpt-4o-transcribeOpenAI
3.17%2.863.53
2.86%
3.53%
96.83%
2.28%
6.21%
3.16%
3.85%
96.15%
Accuracy
Word Error Rate (WER)
3.17%#20/22
WER 95% CI (low)
2.86%#21/22
WER 95% CI (high)
3.53%#20/22
Word Accuracy
96.83%#20/22
Character Error Rate (CER)
2.28%#21/22
Strict WER
6.21%#20/22
Match Error Rate (MER)
3.16%#20/22
Word Information Lost (WIL)
3.85%#20/22
Word Information Preserved (WIP)
96.15%#20/22
Error breakdown
Substitution Rate
0.69%#8/22
Deletion Rate
2.12%#22/22
Insertion Rate
0.36%#1/22
Sentence Error Rate (SER)
51.8%#21/22
Exact Match Rate
48.2%#21/22
Length Ratio
0.982×#22/22
Robustness
Mean WER
8.59%#21/22
Median WER
1.28%#21/22
P90 WER
10.40%#22/22
Severe Error Rate
1.0%#19/22
Empty Output Rate
0.0%#1/22
Success Rate
100.0%#1/22
By clip length
WER - short audio (under 5 s)
5.17%#19/22
WER - medium audio (5 to 15 s)
4.66%#22/22
WER - long audio (over 15 s)
2.20%#21/22
Speed
Median Latency
1.2 s#10/22
P95 Latency
2.0 s#10/22
Real-Time Factor (median)
0.10×#10/22
21
3.35%3.103.63
3.10%
3.63%
96.65%
1.88%
6.37%
3.34%
4.79%
95.21%
22
3.89%2.206.09
2.20%
6.09%
96.11%
3.23%
6.86%
3.81%
4.59%
95.41%
Blue marks the best value in each column; a dash means the metric does not apply to that model. Open a model for all 27 metrics and where it places on each.Showing 22 of 22
Best on each axis

The winners, one question at a time.

Most accurate
1.53%Word Error Rate (WER)qwen3-asr-flash-2026-02-10Qwen
Lowest character error
0.87%Character Error Rate (CER)mai-transcribe-1.5Microsoft
Most exact matches
66.1%Exact Match Rateqwen3-asr-flash-2026-02-10Qwen
Most robust
5.72%P90 WERqwen3-asr-flash-2026-02-10Qwen
Fewest dropped words
0.26%Deletion Ratemai-transcribe-1.5Microsoft
Fewest invented words
0.36%Insertion Rategpt-4o-transcribeOpenAI
Fastest
0.4 sMedian Latencynova-3Deepgram
Glossary

What every column means.

Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.

Accuracy

Word- and character-level agreement with the reference, after normalization — and without it.

Word Error Rate (WER)% · lower is better
Headline metric. (substituted + deleted + inserted words) / words in the ground truth, after normalization. Pooled across all clips.
WER 95% CI (low)% · lower is better
Lower bound of the 95% bootstrap confidence interval for WER.
WER 95% CI (high)% · lower is better
Upper bound of the 95% bootstrap confidence interval for WER.
Word Accuracy% · higher is better
100 − WER (floored at 0).
Character Error Rate (CER)% · lower is better
Same as WER but counted on characters; more forgiving of near-miss spellings.
Strict WER% · lower is better
WER with minimal normalization (only lowercase and punctuation removal). Shows how closely raw output matches the ground-truth style.
Match Error Rate (MER)% · lower is better
Errors / (errors + correct words). Bounded at 100% even with heavy hallucination.
Word Information Lost (WIL)% · lower is better
Share of word-level information lost between ground truth and output.
Word Information Preserved (WIP)% · higher is better
100 − WIL.
Error breakdown

What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.

Substitution Rate% · lower is better
Wrong words, as a share of ground-truth words.
Deletion Rate% · lower is better
Missed words (omissions), as a share of ground-truth words.
Insertion Rate% · lower is better
Extra words not spoken (fabrications / hallucinations), as a share of ground-truth words.
Sentence Error Rate (SER)% · lower is better
Share of clips with at least one word error.
Exact Match Rate% · higher is better
Share of clips transcribed perfectly after normalization.
Length Ratiox · closer to 1 is better
Output word count / ground-truth word count. Above 1 suggests hallucination, below 1 suggests skipped speech.
Robustness

How the error is spread across clips — the typical one, the hardest tenth, the outright failures.

Mean WER% · lower is better
Average of per-clip WER (every clip weighted equally).
Median WER% · lower is better
WER of the typical clip.
P90 WER% · lower is better
WER on the hardest 10% of clips; measures robustness.
Severe Error Rate% · lower is better
Share of clips with WER of 50% or more.
Empty Output Rate% · lower is better
Share of clips where the model returned no words.
Success Rate% · higher is better
Share of requests that returned a transcript.
By clip length

WER split by clip duration. Short clips give a model the least context to recover from.

WER - short audio (under 5 s)% · lower is better
WER on short clips.
WER - medium audio (5 to 15 s)% · lower is better
WER on medium-length clips.
WER - long audio (over 15 s)% · lower is better
WER on long clips.
Speed

Request time and real-time factor. Measured under concurrent load, so indicative only.

Median Latencys · lower is better
Median request time. Measured under concurrent load; indicative only.
P95 Latencys · lower is better
95th percentile request time. Measured under concurrent load; indicative only.
Real-Time Factor (median)x · lower is better
Processing time / audio duration. Measured under concurrent load; indicative only.

Want your model on this board?

We run it on the same audio, under the same scoring, and publish it next to the other 22.