Authenticate and proceed
- Caller
- Customer with a partly forgotten reference number
- Goal
- Verify identity, then continue the original request
Static eval sets cannot interrupt. Our environments are live speech-to-speech calls with trained speakers playing callers who lose the thread, change their mind and read the account number one digit at a time — scored on whether the problem was actually solved.
The scenario families, signals and connection spec below are fixed. We are running them with a small first cohort so the scoring gets argued with by the teams shipping voice agents, before it is sold as a finished product. If your traffic looks like one of these sectors, that is the cohort we want.
Terminals, browsers and repos are cheap to instantiate and free to run a thousand times over. A voice environment needs a person on the other end of the line — which is why the shelf is empty, and why we are the ones filling it.
The environments market grew up around agents that type. Terminals, browsers and repos are cheap to instantiate and free to run a thousand times. A voice environment needs a person on the other end of the line, so almost nobody builds them.
Not in step forty of a plan. It talks over the caller, misses the correction, or answers the question that was abandoned two turns ago. None of that is reachable from a text transcript of a successful call.
The bottleneck is recruiting, training and paying people to play a caller who interrupts on cue, in the right accent, from the right room. That is the operation we already run for recording and grading.
Each sector ships with the same six families, tuned to its own vocabulary and policy. You weight the mix.
Latency is easy to publish and easy to game. We report it, but the environment is scored on task success first.
See the speech-to-speech benchmarkDid the caller's stated goal get met? Scored against a rubric written by an operator from that sector.
End of caller speech to first audible token. Reported as median and p95, not average.
On interruption: stop within 300ms, then act on the new intent rather than the old one.
When the model mishears, does it ask for the one missing token or restart the whole exchange?
Policy held under pressure — no promises the business does not make.
Rated by the same operators who take these calls. Neutral under pressure counts; cheerful does not.
Per-turn transcript, latency trace, rubric hits and misses, and the audio of every failed call.
caller: no no not that card — the other one, ending [overlap] four seven
Every failed call comes back with its audio attached. A score you cannot listen to is a number you cannot act on.
A websocket in, audio frames out. If your model speaks, it can be scored — no SDK to adopt.
Ten sectors, six scenario families each. Weight them the way your traffic actually looks.
Personas are played by trained speakers, including the ones who talk over your agent.
Per-turn transcript, latency trace, rubric hits and misses, and the audio of every failed call.