PERIT-ALIGN · Leaderboard

Word alignment, measured against human experts

14 word-alignment models run on the same real-world audio and scored word by word against timestamps marked and reviewed by people.

RankModelBoundary MAEMedian boundary errorP90 boundary errorStart MAEEnd MAE
14
258.0 ms
240.0 ms
470.0 ms
290.4 ms
225.5 ms
13
116.8 ms
80.0 ms
240.0 ms
85.3 ms
148.4 ms
12
94.7 ms
60.0 ms
200.0 ms
100.8 ms
88.7 ms
11whisper-large-v3-turboOpenAI
97.8 ms
60.0 ms
220.0 ms
110.3 ms
85.4 ms
By tolerance
Timing Accuracy @50 ms
19.63%#10/14
Timing Accuracy @100 ms
46.95%#11/14
Boundaries within 20 ms
19.19%#11/14
Boundaries within 50 ms
43.71%#10/14
Boundaries within 100 ms
68.95%#12/14
Boundaries within 200 ms
88.80%#12/14
Boundary error
Boundary MAE
97.8 ms#12/14
Median boundary error
60.0 ms#10/14
P90 boundary error
220.0 ms#12/14
Start MAE
110.3 ms#13/14
End MAE
85.4 ms#11/14
Drift & shape
Start bias
-94.8 ms#13/14
End bias
-42.7 ms#11/14
Mean IoU
57.36%#10/14
Zero-duration words
0.08%#8/14
Coverage & speed
Word coverage
96.49%#11/14
Transcript WER
4.02%#10/13
Success Rate
100.0%#1/14
Median Latency
1.6 s#10/14
10
66.6 ms
50.0 ms
140.0 ms
84.9 ms
48.3 ms
09
90.1 ms
49.6 ms
169.4 ms
98.6 ms
81.6 ms
08
76.6 ms
80.0 ms
160.0 ms
68.8 ms
84.5 ms
07
68.6 ms
52.0 ms
136.0 ms
70.8 ms
66.4 ms
06
70.0 ms
0.0 ms
180.0 ms
102.1 ms
37.8 ms
05
68.2 ms
0.0 ms
160.0 ms
98.6 ms
37.8 ms
04
52.7 ms
38.0 ms
104.0 ms
57.5 ms
47.9 ms
03
45.6 ms
40.0 ms
80.0 ms
47.3 ms
43.8 ms
02
24.5 ms
0.0 ms
80.0 ms
25.0 ms
23.9 ms
01
23.8 ms
16.0 ms
47.0 ms
22.9 ms
24.6 ms
Blue marks the best value in each column; a dash means the metric does not apply to that model. Open a model for all 19 metrics and where it places on each. Perit's model is a forced aligner: it places the words of a known transcript in time, while the other models transcribe the audio and time the words in one pass.Showing 14 of 14
Best on each axis

The winners, one question at a time.

Most accurate timing
97.85%Timing Accuracy @100 msPerit Golden Data Fine-tuned ModelPerit
Best speech-to-text timestamps
92.33%Timing Accuracy @100 mstranscribe-1Fish Audio
Tightest timing
83.90%Timing Accuracy @50 msPerit Golden Data Fine-tuned ModelPerit
Smallest boundary error
23.8 msBoundary MAEPerit Golden Data Fine-tuned ModelPerit
Best span overlap
84.98%Mean IoUtranscribe-1Fish Audio
Widest coverage
97.70%Word coveragemai-transcribe-2Microsoft
Glossary

What every column means.

Each model's words are matched to the reference words and compared boundary by boundary — start against start, end against end. Coverage and the transcript's own WER are reported next to the timing, so a model cannot look precise by timing only the easy words.

By tolerance

How many words and boundaries land inside each window, from 20 ms to 200 ms.

Timing Accuracy @50 ms% · higher is better
The same, at a 50 ms tolerance — how many words a model places tightly, not just roughly.
Timing Accuracy @100 ms% · higher is better
Headline metric. Share of reference words whose start and end both land within 100 ms of the reference timing.
Boundaries within 20 ms% · higher is better
Share of individual word boundaries (starts and ends counted separately) within 20 ms of the reference.
Boundaries within 50 ms% · higher is better
Share of word boundaries within 50 ms of the reference.
Boundaries within 100 ms% · higher is better
Share of word boundaries within 100 ms of the reference.
Boundaries within 200 ms% · higher is better
Share of word boundaries within 200 ms of the reference.
Boundary error

How far off the boundaries are, in milliseconds — the typical one, the worst tenth, starts against ends.

Boundary MAEms · lower is better
Mean absolute gap between predicted and reference word boundaries, starts and ends pooled. The tie-breaker in the ranking.
Median boundary errorms · lower is better
Boundary error of the typical boundary.
P90 boundary errorms · lower is better
Boundary error at the 90th percentile — how far off the worst tenth of boundaries land.
Start MAEms · lower is better
Mean absolute error of word start times alone.
End MAEms · lower is better
Mean absolute error of word end times alone.
Drift & shape

Whether a model runs systematically early or late, and whether its word spans have real length.

Start biasms · closer to 0 is better
Average signed offset of start times against the reference. Near zero means no systematic drift; the sign shows which way a model leans.
End biasms · closer to 0 is better
Average signed offset of end times against the reference.
Mean IoU% · higher is better
Average overlap between each word's predicted and reference time span (intersection over union).
Zero-duration words% · lower is better
Share of returned words with no length — start and end on the same timestamp.
Coverage & speed

How many words got a timestamp at all, how good the transcript underneath was, and request time — measured under concurrent load, so indicative only.

Word coverage% · higher is better
Share of reference words the model returned a timestamp for. Timing is only scored on words that are covered.
Transcript WER% · lower is better
Word error rate of the transcript the timestamps came with. Context only — it does not affect the ranking, and a forced aligner, which is given the transcript, has none.
Success Rate% · higher is better
Share of requests that returned timestamps.
Median Latencys · lower is better
Median request time. Measured under concurrent load; indicative only.

Want your model on this board?

We run it on the same audio, under the same scoring, and publish it next to the other 14.