PERIT-S2S

The model has to hold a conversation, not finish a sentence.

PERIT-S2S puts voice models into live calls with scripted callers who interrupt, change their mind and give the account number one digit at a time. Scored on whether the caller's problem was actually solved.

Task success (%)

The caller's stated goal was met, verified against a rubric written by an operator from that sector.

Median turn latency (ms, lower is better)

End of caller speech to first audible token of the reply.

Barge-in recovery (%)

Interrupted mid-answer, the model stops, listens and picks up the new intent.

Overall leaderboard

Each model’s mean across all 10 sectors. Open a sector below for the run behind it.

Illustrative figures. Public runs begin when the held-out sets close.
ModelTask success (%)Median turn latencyBarge-in recovery
Human control · sector agentHuman control
92.4 ± 1.7
605ms95.8%
Frontier realtime · lab AFrontier
61.5 ± 1.6
751ms67.8%
Frontier realtime · lab BFrontier
56.9 ± 1.6
884ms61.7%
Cascaded stack (ASR → LLM → TTS)Specialist
51.8 ± 1.5
1444ms44.7%
Frontier multimodal · lab CFrontier
51.0 ± 1.4
1006ms54.3%
Open weights · 8B speechOpen weights
29.7 ± 1.5
1150ms33.5%
Sectors

The same call, ten different rooms.

A refund request in retail and a claim in insurance share a shape. They do not share vocabulary, background noise, or what counts as a wrong answer.

Method

How a run works.

  1. 01Callers are scripted personas played by trained speakers, not synthesised — including the ones who talk over the agent.
  2. 02Every call is scored against the rubric by two reviewers; a third adjudicates splits.
  3. 038 calls per persona per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a licensed agent from that sector handling the same personas blind.