PERIT-S2S · Education

Admissions, fees and course guidance, often with a parent and a student on the same line.

Audio comes from admissions desks, student helplines, parent calls. Scored on task success, with a working education operator as the control.

740
Clips in the target held-out set
58h
Audio, never published
6
Languages in the set
1
Working education operator as control

Leaderboard

Illustrative figures. Public runs begin when the held-out sets close.
ModelTask success (%)Median turn latencyBarge-in recovery
Human control · sector agentHuman control
93.5 ± 1.4
713ms94%
Frontier realtime · lab BFrontier
62.7 ± 1.9
1045ms66.4%
Cascaded stack (ASR → LLM → TTS)Specialist
59.6 ± 1.1
1381ms48.3%
Frontier realtime · lab AFrontier
58.3 ± 1.3
804ms67.3%
Frontier multimodal · lab CFrontier
53.3 ± 1.1
824ms58.8%
Open weights · 8B speechOpen weights
36.3 ± 2.0
902ms40.3%
Task families

What the set actually asks for.

  • Turn-taking & barge-in
  • Task completion
  • Instruction adherence
  • Disambiguation & repair
  • Escalation & refusal
  • Tone under pressure
Method & limits

How a run on this set is produced.

  1. 01Callers are scripted personas played by trained speakers, not synthesised — including the ones who talk over the agent.
  2. 02Every call is scored against the rubric by two reviewers; a third adjudicates splits.
  3. 038 calls per persona per model. Reported intervals are bootstrapped at 95%.
  4. 04Human control is a licensed agent from that sector handling the same personas blind.

Limits, stated up front: 6 languages only, single-channel audio, and no cross-sector transfer. The control is one operator per sector, so the human line carries its own error bar — it is a working standard, not a ceiling.