PERIT-ALIGN · Leaderboard

Word alignment, measured against human experts

14 word-alignment models run on the same real-world audio and scored word by word against timestamps marked and reviewed by people.

RankModelTiming Accuracy @100 msTiming Accuracy @50 msBoundary MAEP90 boundary errorMean IoUWord coverage
01
97.85%
83.90%
23.8 ms
47.0 ms
81.85%
100.00%
02
92.33%
67.69%
24.5 ms
80.0 ms
84.98%
97.58%
03
85.04%
43.96%
45.6 ms
80.0 ms
69.52%
96.68%
04
78.04%
38.89%
52.7 ms
104.0 ms
64.58%
97.56%
05
77.41%
59.52%
68.2 ms
160.0 ms
72.03%
97.24%
06
76.60%
58.75%
70.0 ms
180.0 ms
71.01%
97.56%
07
63.45%
22.77%
68.6 ms
136.0 ms
53.75%
97.64%
08
60.19%
13.09%
76.6 ms
160.0 ms
59.21%
96.91%
09whisper-large-v3OpenAI
59.89%
25.79%
90.1 ms
169.4 ms
58.36%
95.84%
By tolerance
Timing Accuracy @50 ms
25.79%#7/14
Timing Accuracy @100 ms
59.89%#9/14
Boundaries within 20 ms
21.70%#8/14
Boundaries within 50 ms
50.59%#8/14
Boundaries within 100 ms
77.61%#10/14
Boundaries within 200 ms
92.05%#8/14
Boundary error
Boundary MAE
90.1 ms#10/14
Median boundary error
49.6 ms#7/14
P90 boundary error
169.4 ms#9/14
Start MAE
98.6 ms#9/14
End MAE
81.6 ms#9/14
Drift & shape
Start bias
-10.4 ms#3/14
End bias
-15.7 ms#6/14
Mean IoU
58.36%#8/14
Zero-duration words
0.01%#7/14
Coverage & speed
Word coverage
95.84%#13/14
Transcript WER
4.14%#12/13
Success Rate
100.0%#1/14
Median Latency
1.2 s#8/14
10
59.85%
25.60%
66.6 ms
140.0 ms
57.86%
97.70%
11
46.95%
19.63%
97.8 ms
220.0 ms
57.36%
96.49%
12
45.86%
15.73%
94.7 ms
200.0 ms
54.14%
96.37%
13
30.96%
4.16%
116.8 ms
240.0 ms
27.89%
97.67%
14
5.77%
0.65%
258.0 ms
470.0 ms
14.04%
94.96%
Blue marks the best value in each column; a dash means the metric does not apply to that model. Open a model for all 19 metrics and where it places on each. Perit's model is a forced aligner: it places the words of a known transcript in time, while the other models transcribe the audio and time the words in one pass.Showing 14 of 14
Best on each axis

The winners, one question at a time.

Most accurate timing
97.85%Timing Accuracy @100 msPerit Golden Data Fine-tuned ModelPerit
Best speech-to-text timestamps
92.33%Timing Accuracy @100 mstranscribe-1Fish Audio
Tightest timing
83.90%Timing Accuracy @50 msPerit Golden Data Fine-tuned ModelPerit
Smallest boundary error
23.8 msBoundary MAEPerit Golden Data Fine-tuned ModelPerit
Best span overlap
84.98%Mean IoUtranscribe-1Fish Audio
Widest coverage
97.70%Word coveragemai-transcribe-2Microsoft
Glossary

What every column means.

Each model's words are matched to the reference words and compared boundary by boundary — start against start, end against end. Coverage and the transcript's own WER are reported next to the timing, so a model cannot look precise by timing only the easy words.

By tolerance

How many words and boundaries land inside each window, from 20 ms to 200 ms.

Timing Accuracy @50 ms% · higher is better
The same, at a 50 ms tolerance — how many words a model places tightly, not just roughly.
Timing Accuracy @100 ms% · higher is better
Headline metric. Share of reference words whose start and end both land within 100 ms of the reference timing.
Boundaries within 20 ms% · higher is better
Share of individual word boundaries (starts and ends counted separately) within 20 ms of the reference.
Boundaries within 50 ms% · higher is better
Share of word boundaries within 50 ms of the reference.
Boundaries within 100 ms% · higher is better
Share of word boundaries within 100 ms of the reference.
Boundaries within 200 ms% · higher is better
Share of word boundaries within 200 ms of the reference.
Boundary error

How far off the boundaries are, in milliseconds — the typical one, the worst tenth, starts against ends.

Boundary MAEms · lower is better
Mean absolute gap between predicted and reference word boundaries, starts and ends pooled. The tie-breaker in the ranking.
Median boundary errorms · lower is better
Boundary error of the typical boundary.
P90 boundary errorms · lower is better
Boundary error at the 90th percentile — how far off the worst tenth of boundaries land.
Start MAEms · lower is better
Mean absolute error of word start times alone.
End MAEms · lower is better
Mean absolute error of word end times alone.
Drift & shape

Whether a model runs systematically early or late, and whether its word spans have real length.

Start biasms · closer to 0 is better
Average signed offset of start times against the reference. Near zero means no systematic drift; the sign shows which way a model leans.
End biasms · closer to 0 is better
Average signed offset of end times against the reference.
Mean IoU% · higher is better
Average overlap between each word's predicted and reference time span (intersection over union).
Zero-duration words% · lower is better
Share of returned words with no length — start and end on the same timestamp.
Coverage & speed

How many words got a timestamp at all, how good the transcript underneath was, and request time — measured under concurrent load, so indicative only.

Word coverage% · higher is better
Share of reference words the model returned a timestamp for. Timing is only scored on words that are covered.
Transcript WER% · lower is better
Word error rate of the transcript the timestamps came with. Context only — it does not affect the ranking, and a forced aligner, which is given the transcript, has none.
Success Rate% · higher is better
Share of requests that returned timestamps.
Median Latencys · lower is better
Median request time. Measured under concurrent load; indicative only.

Want your model on this board?

We run it on the same audio, under the same scoring, and publish it next to the other 14.