{"id":687,"date":"2026-09-16T08:33:31","date_gmt":"2026-09-16T08:33:31","guid":{"rendered":"https:\/\/perit.ai\/blogs\/how-we-benchmark-speech-to-speech-models\/"},"modified":"2026-09-16T13:40:52","modified_gmt":"2026-09-16T13:40:52","slug":"how-we-benchmark-speech-to-speech-models","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/how-we-benchmark-speech-to-speech-models\/","title":{"rendered":"How We Benchmark Speech-to-Speech Models"},"content":{"rendered":"<p>Word error rate is easy to compute and, on its own, a poor predictor of whether a voice model is actually pleasant to talk to.<\/p>\n<h2>Quick comparison: what gets measured<\/h2>\n<p>Our benchmark scores four axes: intelligibility, latency, prosody\/naturalness, and turn-taking behavior \u2014 how gracefully a model handles interruption and overlap.<\/p>\n<h3>Intelligibility and latency<\/h3>\n<p>These are the table-stakes metrics \u2014 if a response is slow or hard to understand, nothing else matters. We measure both under realistic network conditions, not a clean lab environment.<\/p>\n<h3>Naturalness<\/h3>\n<p>Scored by human raters blind to which model produced which clip, on a rubric built from the same grading process used for training data itself.<\/p>\n<h2>Why this is harder than text benchmarks<\/h2>\n<p>Text benchmarks can be scored automatically against a reference answer. Voice interaction quality is inherently subjective, which is why human grading \u2014 not just automated scoring \u2014 sits at the center of the methodology.<\/p>\n<h2>Frequently asked questions<\/h2>\n<p><strong>Is the benchmark public?<\/strong> A summary leaderboard is published on the Research page; the full clip set is available under a research license.<\/p>\n<p><strong>How often is it updated?<\/strong> Quarterly, or sooner when a new class of model warrants a fresh comparison round.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A look at the methodology behind Perit AI&#8217;s speech-to-speech benchmark, and why naturalness is harder to score than accuracy.<\/p>\n","protected":false},"author":1,"featured_media":724,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[19],"tags":[],"class_list":["post-687","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-research"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/687","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=687"}],"version-history":[{"count":1,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/687\/revisions"}],"predecessor-version":[{"id":725,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/687\/revisions\/725"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/724"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=687"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=687"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=687"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}