PERIT-S2S · Healthcare

Appointment intake, triage lines and medication questions, dense with drug names and dosages.

Audio comes from front-desk intake, nurse lines, pharmacy callbacks. Scored on task success, with a working healthcare operator as the control.

1,020
Clips in the target held-out set
81h
Audio, never published
5
Languages in the set
1
Working healthcare operator as control

Leaderboard

Illustrative figures. Public runs begin when the held-out sets close.
ModelTask success (%)Median turn latencyBarge-in recovery
Human control · sector agentHuman control
92.5 ± 0.9
696ms93.4%
Cascaded stack (ASR → LLM → TTS)Specialist
57.0 ± 1.8
1648ms45.9%
Frontier realtime · lab AFrontier
55.3 ± 1.0
705ms67.8%
Frontier realtime · lab BFrontier
53.4 ± 2.2
1041ms57.7%
Frontier multimodal · lab CFrontier
48.5 ± 2.0
956ms51.8%
Open weights · 8B speechOpen weights
26.1 ± 2.3
1192ms34.4%
Task families

What the set actually asks for.

  • Turn-taking & barge-in
  • Task completion
  • Instruction adherence
  • Disambiguation & repair
  • Escalation & refusal
  • Tone under pressure
Method & limits

How a run on this set is produced.

  1. 01Callers are scripted personas played by trained speakers, not synthesised — including the ones who talk over the agent.
  2. 02Every call is scored against the rubric by two reviewers; a third adjudicates splits.
  3. 038 calls per persona per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a licensed agent from that sector handling the same personas blind.

Limits, stated up front: 5 languages only, single-channel audio, and no cross-sector transfer. The control is one operator per sector, so the human line carries its own error bar — it is a working standard, not a ceiling.