How We Benchmark Speech-to-Speech Models

How We Benchmark Speech-to-Speech Models

Word error rate is easy to compute and, on its own, a poor predictor of whether a voice model is actually pleasant to talk to.

Quick comparison: what gets measured

Our benchmark scores four axes: intelligibility, latency, prosody/naturalness, and turn-taking behavior — how gracefully a model handles interruption and overlap.

Intelligibility and latency

These are the table-stakes metrics — if a response is slow or hard to understand, nothing else matters. We measure both under realistic network conditions, not a clean lab environment.

Naturalness

Scored by human raters blind to which model produced which clip, on a rubric built from the same grading process used for training data itself.

Why this is harder than text benchmarks

Text benchmarks can be scored automatically against a reference answer. Voice interaction quality is inherently subjective, which is why human grading — not just automated scoring — sits at the center of the methodology.

Frequently asked questions

Is the benchmark public? A summary leaderboard is published on the Research page; the full clip set is available under a research license.

How often is it updated? Quarterly, or sooner when a new class of model warrants a fresh comparison round.

Was this article helpful? No ratings yet