22 speech-to-text models run on the same real-world audio and scored word by word against transcripts written and reviewed by people.
Both the reference and the model output go through the same normalization before a single word is counted, so a model is not penalised for writing “2” where the transcriber wrote “two”.
Word- and character-level agreement with the reference, after normalization — and without it.
What kind of mistake a model makes: the wrong word, a dropped word, or an invented one.
How the error is spread across clips — the typical one, the hardest tenth, the outright failures.
WER split by clip duration. Short clips give a model the least context to recover from.
Request time and real-time factor. Measured under concurrent load, so indicative only.
We run it on the same audio, under the same scoring, and publish it next to the other 22.