Word error rate is easy to compute and, on its own, a poor predictor of whether a voice model is actually pleasant to talk to.
Quick comparison: what gets measured
Our benchmark scores four axes: intelligibility, latency, prosody/naturalness, and turn-taking behavior — how gracefully a model handles interruption and overlap.
Intelligibility and latency
These are the table-stakes metrics — if a response is slow or hard to understand, nothing else matters. We measure both under realistic network conditions, not a clean lab environment.
Naturalness
Scored by human raters blind to which model produced which clip, on a rubric built from the same grading process used for training data itself.
Why this is harder than text benchmarks
Text benchmarks can be scored automatically against a reference answer. Voice interaction quality is inherently subjective, which is why human grading — not just automated scoring — sits at the center of the methodology.
Frequently asked questions
Is the benchmark public? A summary leaderboard is published on the Research page; the full clip set is available under a research license.
How often is it updated? Quarterly, or sooner when a new class of model warrants a fresh comparison round.