PERIT-ALIGN scores the word-level timestamps a speech model returns — where each word starts and where it ends — against reference timings, so captions, redaction and audio search land on the right moment.
Accuracy is only one of the questions, so each of these is the best model on its own axis.
Higher is better. A word counts when both its start and its end land within 100 ms of the reference.
Timing accuracy: the share of words placed within 100 ms of the reference. Longer bar is better. Perit's model is a forced aligner: it places the words of a known transcript in time, while the other models transcribe the audio and time the words in one pass.
Each model's words are matched to the reference words and compared boundary by boundary — start against start, end against end. Coverage and the transcript's own WER are reported next to the timing, so a model cannot look precise by timing only the easy words.