PERIT-STT · Leaderboard

Speech recognition, measured against human experts

Ranked by word error rate. Switch tabs to see accuracy, the kind of error, robustness on hard clips, clip length, and what each model costs to run.

RankModelWord Error RateCharacter Error RateWord AccuracyExact Match RatePrice per 1,000 minMedian Latency
01
0.91%0.561.33
0.50%
99.09%
78.0%
$2.03
9.3 s
02
1.28%0.791.83
0.74%
98.72%
70.0%
$6.27
7.1 s
03
1.33%0.881.84
0.63%
98.67%
71.0%
$6.27
7.5 s
04
1.36%0.901.94
0.64%
98.64%
71.0%
$0.20
6.7 s
05qwen3-asr-1.7bQwen
1.36%0.871.86
0.75%
98.64%
70.0%
$0.45
5.8 s
Accuracy
Word Error Rate (WER)
1.36%#4/22
WER 95% CI (low)
0.87%#3/22
WER 95% CI (high)
1.86%#4/22
Word Accuracy
98.64%#4/22
Character Error Rate (CER)
0.75%#6/22
Strict WER
5.06%#7/22
Match Error Rate (MER)
1.36%#4/22
Word Information Lost (WIL)
1.92%#4/22
Word Information Preserved (WIP)
98.08%#4/22
Error breakdown
Substitution Rate
0.57%#4/22
Deletion Rate
0.51%#8/22
Insertion Rate
0.28%#12/22
Sentence Error Rate (SER)
30.0%#4/22
Exact Match Rate
70.0%#4/22
Length Ratio
0.998×#5/22
Robustness
Mean WER
3.68%#16/22
Median WER
0.00%#1/22
P90 WER
5.73%#7/22
Severe Error Rate
2.0%#18/22
Empty Output Rate
0.0%#1/22
Success Rate
100.0%#1/22
By clip length
WER - short audio (under 5 s)
2.03%#5/22
WER - medium audio (5 to 15 s)
1.55%#3/22
WER - long audio (over 15 s)
1.13%#7/22
Cost & speed
Price per 1,000 min
$0.45#4/22
Total Cost
$0.0093#4/22
Median Latency
5.8 s#1/22
P95 Latency
67.4 s#13/22
Real-Time Factor (median)
0.69×#3/22
06
1.45%1.011.94
0.78%
98.55%
65.0%
$16.72
9.9 s
07
1.45%0.991.94
0.80%
98.55%
69.0%
$1.74
7.8 s
08
1.48%1.011.94
0.70%
98.52%
67.0%
$3.75
10.3 s
09
1.48%1.002.11
0.75%
98.52%
66.0%
$3.01
10.0 s
10
1.59%1.072.15
0.91%
98.41%
68.0%
$4.70
8.3 s
11
1.65%1.192.20
0.83%
98.35%
62.0%
$1.67
7.5 s
12
1.82%1.222.42
1.06%
98.18%
66.0%
$1.75
10.5 s
13
1.96%1.422.52
1.07%
98.04%
55.0%
$2.89
8.6 s
14
1.99%1.452.55
1.03%
98.01%
61.0%
$1.50
7.0 s
15
2.06%1.552.64
1.19%
97.94%
51.5%
$4.30
8.0 s
16
2.13%1.552.76
1.22%
97.87%
58.0%
$6.27
8.3 s
17
2.24%1.563.04
1.43%
97.76%
57.0%
$1.00
7.3 s
18
2.27%1.603.08
1.32%
97.73%
53.0%
$3.00
9.4 s
19
2.33%1.763.04
1.42%
97.67%
49.0%
$0.23
8.5 s
20
2.35%1.753.08
1.49%
97.65%
50.0%
$1.59
7.6 s
21
3.12%2.314.08
1.59%
96.88%
52.0%
$0.20
10.0 s
22
3.94%2.645.75
2.92%
96.06%
47.0%
$3.44
6.8 s
Blue marks the best value in each column. Open a model for all 29 metrics and where it places on each.Showing 22 of 22
Best on each axis

The winners, one question at a time.

Most accurate
0.91%Word Error Rate (WER)qwen3-asr-flash-2026-02-10Qwen
Lowest character error
0.50%Character Error Rate (CER)qwen3-asr-flash-2026-02-10Qwen
Most exact matches
78.0%Exact Match Rateqwen3-asr-flash-2026-02-10Qwen
Most robust
4.08%P90 WERqwen3-asr-flash-2026-02-10Qwen
Fewest dropped words
0.34%Deletion Ratemai-transcribe-1.5Microsoft
Fewest invented words
0.17%Insertion Ratevoxtral-small-24b-2507-sttMistral AI
Lowest price
$0.20Price per 1,000 minnemotron-3.5-asr-streaming-multilingual-0.6bNVIDIA · tied with qwen3-asr-0.6b
Fastest
5.8 sMedian Latencyqwen3-asr-1.7bQwen
Glossary

What every column means.

Rates are pooled across all clips unless the name says mean, median or P90. Every metric is computed after the same normalization on both the reference and the model output.

Accuracy

Word- and character-level agreement with the reference, after normalization — and without it.

Word Error Rate (WER)% · lower is better
Headline metric. (substituted + deleted + inserted words) / words in the ground truth, after normalization. Pooled across all clips.
WER 95% CI (low)% · lower is better
Lower bound of the 95% bootstrap confidence interval for WER.
WER 95% CI (high)% · lower is better
Upper bound of the 95% bootstrap confidence interval for WER.
Word Accuracy% · higher is better
100 − WER (floored at 0).
Character Error Rate (CER)% · lower is better
Same as WER but counted on characters; more forgiving of near-miss spellings.
Strict WER% · lower is better
WER with minimal normalization (only lowercase and punctuation removal). Shows how closely raw output matches the ground-truth style.
Match Error Rate (MER)% · lower is better
Errors / (errors + correct words). Bounded at 100% even with heavy hallucination.
Word Information Lost (WIL)% · lower is better
Share of word-level information lost between ground truth and output.
Word Information Preserved (WIP)% · higher is better
100 − WIL.
Error breakdown

What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.

Substitution Rate% · lower is better
Wrong words, as a share of ground-truth words.
Deletion Rate% · lower is better
Missed words (omissions), as a share of ground-truth words.
Insertion Rate% · lower is better
Extra words not spoken (fabrications / hallucinations), as a share of ground-truth words.
Sentence Error Rate (SER)% · lower is better
Share of clips with at least one word error.
Exact Match Rate% · higher is better
Share of clips transcribed perfectly after normalization.
Length Ratiox · closer to 1 is better
Output word count / ground-truth word count. Above 1 suggests hallucination, below 1 suggests skipped speech.
Robustness

How the error is spread across clips — the typical one, the hardest tenth, the outright failures.

Mean WER% · lower is better
Average of per-clip WER (every clip weighted equally).
Median WER% · lower is better
WER of the typical clip.
P90 WER% · lower is better
WER on the hardest 10% of clips; measures robustness.
Severe Error Rate% · lower is better
Share of clips with WER of 50% or more.
Empty Output Rate% · lower is better
Share of clips where the model returned no words.
Success Rate% · higher is better
Share of requests that returned a transcript.
By clip length

WER split by clip duration. Short clips give a model the least context to recover from.

WER - short audio (under 5 s)% · lower is better
WER on short clips.
WER - medium audio (5 to 15 s)% · lower is better
WER on medium-length clips.
WER - long audio (over 15 s)% · lower is better
WER on long clips.
Cost & speed

Billed cost and request time. Latency was measured under concurrent load and is indicative only.

Price per 1,000 minUSD · lower is better
Effective cost to transcribe 1,000 minutes of audio, from actual billing.
Total CostUSD · lower is better
Actual billed cost for the full benchmark run.
Median Latencys · lower is better
Median request time. Measured under concurrent load; indicative only.
P95 Latencys · lower is better
95th percentile request time. Measured under concurrent load; indicative only.
Real-Time Factor (median)x · lower is better
Processing time / audio duration. Measured under concurrent load; indicative only.

Want your model on this board?

We run it on the same audio, through the same normalization, and publish it next to the other 22.