PERIT-S2S · Insurance

First-notice-of-loss calls taken from roadsides and hospital corridors, rarely from a quiet room.

Audio comes from fnol intake, claims status, renewals. Scored on task success, with a working insurance operator as the control.

960
Clips in the target held-out set
74h
Audio, never published
5
Languages in the set
1
Working insurance operator as control

Leaderboard

Illustrative figures. Public runs begin when the held-out sets close.
ModelTask success (%)Median turn latencyBarge-in recovery
Human control · sector agentHuman control
90.1 ± 1.8
503ms97.8%
Frontier realtime · lab AFrontier
68.0 ± 1.6
660ms64.4%
Frontier realtime · lab BFrontier
59.8 ± 2.0
828ms61.8%
Frontier multimodal · lab CFrontier
54.7 ± 1.5
1078ms54.1%
Cascaded stack (ASR → LLM → TTS)Specialist
52.6 ± 1.3
1320ms39.5%
Open weights · 8B speechOpen weights
35.8 ± 1.6
1018ms37.3%
Task families

What the set actually asks for.

  • Turn-taking & barge-in
  • Task completion
  • Instruction adherence
  • Disambiguation & repair
  • Escalation & refusal
  • Tone under pressure
Method & limits

How a run on this set is produced.

  1. 01Callers are scripted personas played by trained speakers, not synthesised — including the ones who talk over the agent.
  2. 02Every call is scored against the rubric by two reviewers; a third adjudicates splits.
  3. 038 calls per persona per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a licensed agent from that sector handling the same personas blind.

Limits, stated up front: 5 languages only, single-channel audio, and no cross-sector transfer. The control is one operator per sector, so the human line carries its own error bar — it is a working standard, not a ceiling.