Latency is the easiest speech number to publish and the easiest to game. It belongs in the report, not at the top of it.
A voice agent can answer in 300 milliseconds and still lose the call. It can take a second and a half and solve the problem. If the headline metric is time-to-first-token, teams optimise for the number rather than the outcome, and the benchmark starts rewarding models that interrupt to look fast.
So PERIT-S2S is scored on whether the caller's stated goal was met, against a rubric written by someone who takes those calls. Median and p95 latency sit alongside it, because a two-second gap is a real product problem — just not the definition of a good call.