PERIT-ALIGN · Leaderboard

Word alignment, measured against human experts

14 word-alignment models run on the same real-world audio and scored word by word against timestamps marked and reviewed by people.

RankModelTiming Accuracy @50 msTiming Accuracy @100 msBoundaries within 20 msBoundaries within 50 msBoundaries within 100 msBoundaries within 200 ms
14
0.65%
5.77%
5.27%
6.73%
20.03%
41.00%
13
4.16%
30.96%
18.69%
24.73%
57.19%
83.46%
12
15.73%
45.86%
18.32%
42.24%
69.97%
89.80%
11
19.63%
46.95%
19.19%
43.71%
68.95%
88.80%
10
25.60%
59.85%
19.76%
50.81%
77.62%
96.98%
09
25.79%
59.89%
21.70%
50.59%
77.61%
92.05%
08
13.09%
60.19%
26.54%
36.35%
78.45%
95.74%
07
22.77%
63.45%
20.03%
48.70%
80.23%
96.29%
06qwen3-asr-1.7bQwen
58.75%
76.60%
64.79%
75.52%
87.56%
90.15%
By tolerance
Timing Accuracy @50 ms
58.75%#4/14
Timing Accuracy @100 ms
76.60%#6/14
Boundaries within 20 ms
64.79%#3/14
Boundaries within 50 ms
75.52%#4/14
Boundaries within 100 ms
87.56%#6/14
Boundaries within 200 ms
90.15%#10/14
Boundary error
Boundary MAE
70.0 ms#8/14
Median boundary error
0.0 ms#1/14
P90 boundary error
180.0 ms#10/14
Start MAE
102.1 ms#12/14
End MAE
37.8 ms#3/14
Drift & shape
Start bias
79.6 ms#12/14
End bias
22.9 ms#8/14
Mean IoU
71.01%#4/14
Zero-duration words
12.47%#13/14
Coverage & speed
Word coverage
97.56%#6/14
Transcript WER
2.99%#5/13
Success Rate
99.8%#13/14
Median Latency
1.7 s#12/14
05
59.52%
77.41%
65.40%
76.17%
88.11%
90.56%
04
38.89%
78.04%
28.02%
62.33%
88.92%
97.81%
03
43.96%
85.04%
28.55%
67.10%
93.50%
98.59%
02
67.69%
92.33%
71.13%
82.82%
96.93%
98.71%
01
83.90%
97.85%
56.59%
91.41%
98.86%
99.73%
Blue marks the best value in each column; a dash means the metric does not apply to that model. Open a model for all 19 metrics and where it places on each. Perit's model is a forced aligner: it places the words of a known transcript in time, while the other models transcribe the audio and time the words in one pass.Showing 14 of 14
Best on each axis

The winners, one question at a time.

Most accurate timing
97.85%Timing Accuracy @100 msPerit Golden Data Fine-tuned ModelPerit
Best speech-to-text timestamps
92.33%Timing Accuracy @100 mstranscribe-1Fish Audio
Tightest timing
83.90%Timing Accuracy @50 msPerit Golden Data Fine-tuned ModelPerit
Smallest boundary error
23.8 msBoundary MAEPerit Golden Data Fine-tuned ModelPerit
Best span overlap
84.98%Mean IoUtranscribe-1Fish Audio
Widest coverage
97.70%Word coveragemai-transcribe-2Microsoft
Glossary

What every column means.

Each model's words are matched to the reference words and compared boundary by boundary — start against start, end against end. Coverage and the transcript's own WER are reported next to the timing, so a model cannot look precise by timing only the easy words.

By tolerance

How many words and boundaries land inside each window, from 20 ms to 200 ms.

Timing Accuracy @50 ms% · higher is better
The same, at a 50 ms tolerance — how many words a model places tightly, not just roughly.
Timing Accuracy @100 ms% · higher is better
Headline metric. Share of reference words whose start and end both land within 100 ms of the reference timing.
Boundaries within 20 ms% · higher is better
Share of individual word boundaries (starts and ends counted separately) within 20 ms of the reference.
Boundaries within 50 ms% · higher is better
Share of word boundaries within 50 ms of the reference.
Boundaries within 100 ms% · higher is better
Share of word boundaries within 100 ms of the reference.
Boundaries within 200 ms% · higher is better
Share of word boundaries within 200 ms of the reference.
Boundary error

How far off the boundaries are, in milliseconds — the typical one, the worst tenth, starts against ends.

Boundary MAEms · lower is better
Mean absolute gap between predicted and reference word boundaries, starts and ends pooled. The tie-breaker in the ranking.
Median boundary errorms · lower is better
Boundary error of the typical boundary.
P90 boundary errorms · lower is better
Boundary error at the 90th percentile — how far off the worst tenth of boundaries land.
Start MAEms · lower is better
Mean absolute error of word start times alone.
End MAEms · lower is better
Mean absolute error of word end times alone.
Drift & shape

Whether a model runs systematically early or late, and whether its word spans have real length.

Start biasms · closer to 0 is better
Average signed offset of start times against the reference. Near zero means no systematic drift; the sign shows which way a model leans.
End biasms · closer to 0 is better
Average signed offset of end times against the reference.
Mean IoU% · higher is better
Average overlap between each word's predicted and reference time span (intersection over union).
Zero-duration words% · lower is better
Share of returned words with no length — start and end on the same timestamp.
Coverage & speed

How many words got a timestamp at all, how good the transcript underneath was, and request time — measured under concurrent load, so indicative only.

Word coverage% · higher is better
Share of reference words the model returned a timestamp for. Timing is only scored on words that are covered.
Transcript WER% · lower is better
Word error rate of the transcript the timestamps came with. Context only — it does not affect the ranking, and a forced aligner, which is given the transcript, has none.
Success Rate% · higher is better
Share of requests that returned timestamps.
Median Latencys · lower is better
Median request time. Measured under concurrent load; indicative only.

Want your model on this board?

We run it on the same audio, under the same scoring, and publish it next to the other 14.