PERIT-ALIGN · Leaderboard

Word alignment, measured against human experts

14 word-alignment models run on the same real-world audio and scored word by word against timestamps marked and reviewed by people.

RankModelWord coverageTranscript WERSuccess RateMedian Latency
01
100.00%
—
100.0%
0.4 s
02
97.58%
3.02%
100.0%
0.6 s
03
96.68%
3.64%
100.0%
1.5 s
04
97.56%
2.82%
100.0%
1.0 s
05
97.24%
3.25%
100.0%
2.1 s
06
97.56%
2.99%
99.8%
1.7 s
07
97.64%
2.63%
100.0%
0.6 s
08
96.91%
3.45%
99.0%
0.5 s
09whisper-large-v3OpenAI
95.84%
4.14%
100.0%
1.2 s
By tolerance
Timing Accuracy @50 ms
25.79%#7/14
Timing Accuracy @100 ms
59.89%#9/14
Boundaries within 20 ms
21.70%#8/14
Boundaries within 50 ms
50.59%#8/14
Boundaries within 100 ms
77.61%#10/14
Boundaries within 200 ms
92.05%#8/14
Boundary error
Boundary MAE
90.1 ms#10/14
Median boundary error
49.6 ms#7/14
P90 boundary error
169.4 ms#9/14
Start MAE
98.6 ms#9/14
End MAE
81.6 ms#9/14
Drift & shape
Start bias
-10.4 ms#3/14
End bias
-15.7 ms#6/14
Mean IoU
58.36%#8/14
Zero-duration words
0.01%#7/14
Coverage & speed
Word coverage
95.84%#13/14
Transcript WER
4.14%#12/13
Success Rate
100.0%#1/14
Median Latency
1.2 s#8/14
10
97.70%
2.64%
100.0%
0.7 s
11
96.49%
4.02%
100.0%
1.6 s
12
96.37%
4.07%
100.0%
1.7 s
13
97.67%
2.96%
100.0%
0.7 s
14
94.96%
5.36%
99.8%
2.4 s
Blue marks the best value in each column; a dash means the metric does not apply to that model. Open a model for all 19 metrics and where it places on each. Perit's model is a forced aligner: it places the words of a known transcript in time, while the other models transcribe the audio and time the words in one pass.Showing 14 of 14
Best on each axis

The winners, one question at a time.

Most accurate timing
97.85%Timing Accuracy @100 msPerit Golden Data Fine-tuned ModelPerit
Best speech-to-text timestamps
92.33%Timing Accuracy @100 mstranscribe-1Fish Audio
Tightest timing
83.90%Timing Accuracy @50 msPerit Golden Data Fine-tuned ModelPerit
Smallest boundary error
23.8 msBoundary MAEPerit Golden Data Fine-tuned ModelPerit
Best span overlap
84.98%Mean IoUtranscribe-1Fish Audio
Widest coverage
97.70%Word coveragemai-transcribe-2Microsoft
Glossary

What every column means.

Each model's words are matched to the reference words and compared boundary by boundary — start against start, end against end. Coverage and the transcript's own WER are reported next to the timing, so a model cannot look precise by timing only the easy words.

By tolerance

How many words and boundaries land inside each window, from 20 ms to 200 ms.

Timing Accuracy @50 ms% · higher is better
The same, at a 50 ms tolerance — how many words a model places tightly, not just roughly.
Timing Accuracy @100 ms% · higher is better
Headline metric. Share of reference words whose start and end both land within 100 ms of the reference timing.
Boundaries within 20 ms% · higher is better
Share of individual word boundaries (starts and ends counted separately) within 20 ms of the reference.
Boundaries within 50 ms% · higher is better
Share of word boundaries within 50 ms of the reference.
Boundaries within 100 ms% · higher is better
Share of word boundaries within 100 ms of the reference.
Boundaries within 200 ms% · higher is better
Share of word boundaries within 200 ms of the reference.
Boundary error

How far off the boundaries are, in milliseconds — the typical one, the worst tenth, starts against ends.

Boundary MAEms · lower is better
Mean absolute gap between predicted and reference word boundaries, starts and ends pooled. The tie-breaker in the ranking.
Median boundary errorms · lower is better
Boundary error of the typical boundary.
P90 boundary errorms · lower is better
Boundary error at the 90th percentile — how far off the worst tenth of boundaries land.
Start MAEms · lower is better
Mean absolute error of word start times alone.
End MAEms · lower is better
Mean absolute error of word end times alone.
Drift & shape

Whether a model runs systematically early or late, and whether its word spans have real length.

Start biasms · closer to 0 is better
Average signed offset of start times against the reference. Near zero means no systematic drift; the sign shows which way a model leans.
End biasms · closer to 0 is better
Average signed offset of end times against the reference.
Mean IoU% · higher is better
Average overlap between each word's predicted and reference time span (intersection over union).
Zero-duration words% · lower is better
Share of returned words with no length — start and end on the same timestamp.
Coverage & speed

How many words got a timestamp at all, how good the transcript underneath was, and request time — measured under concurrent load, so indicative only.

Word coverage% · higher is better
Share of reference words the model returned a timestamp for. Timing is only scored on words that are covered.
Transcript WER% · lower is better
Word error rate of the transcript the timestamps came with. Context only — it does not affect the ranking, and a forced aligner, which is given the transcript, has none.
Success Rate% · higher is better
Share of requests that returned timestamps.
Median Latencys · lower is better
Median request time. Measured under concurrent load; indicative only.

Want your model on this board?

We run it on the same audio, under the same scoring, and publish it next to the other 14.