Environments

Put your voice model on a call with someone who talks over it.

Static eval sets cannot interrupt. Our environments are live speech-to-speech calls with trained speakers playing callers who lose the thread, change their mind and read the account number one digit at a time — scored on whether the problem was actually solved.

6
Scenario families per sector
10
Sectors available
300ms
Barge-in detection window scored
8
Runs per persona, bootstrapped CI
Design partners, first cohort

The scenario families, signals and connection spec below are fixed. We are running them with a small first cohort so the scoring gets argued with by the teams shipping voice agents, before it is sold as a finished product. If your traffic looks like one of these sectors, that is the cohort we want.

Why speech

The environments market grew up around agents that type.

Terminals, browsers and repos are cheap to instantiate and free to run a thousand times over. A voice environment needs a person on the other end of the line — which is why the shelf is empty, and why we are the ones filling it.

Everyone else is building coding environments

The environments market grew up around agents that type. Terminals, browsers and repos are cheap to instantiate and free to run a thousand times. A voice environment needs a person on the other end of the line, so almost nobody builds them.

A voice agent fails in the first ten seconds

Not in step forty of a plan. It talks over the caller, misses the correction, or answers the question that was abandoned two turns ago. None of that is reachable from a text transcript of a successful call.

The hard part is the caller, and the caller is our day job

The bottleneck is recruiting, training and paying people to play a caller who interrupts on cue, in the right accent, from the right room. That is the operation we already run for recording and grading.

Scenarios

Six ways a call goes wrong.

Each sector ships with the same six families, tuned to its own vocabulary and policy. You weight the mix.

Authenticate and proceed

Caller
Customer with a partly forgotten reference number
Goal
Verify identity, then continue the original request
Digits out of order, with corrections

Interrupted mid-answer

Caller
Impatient caller who already knows half the answer
Goal
Stop, absorb the new intent, finish the right task
Barge-in at 1.5s, twice per call

Change of mind

Caller
Caller who reverses the request halfway
Goal
Discard the abandoned intent without re-asking everything
Contradictory instructions, one turn

Phone handed over

Caller
Second speaker takes the call mid-way
Goal
Re-establish who is speaking and what they are allowed to do
New voice, no announcement

Bad line

Caller
Caller from a car, a street, a warehouse
Goal
Complete the task or ask for the one thing it needs repeated
Packet loss, wind, crosstalk

Escalation and refusal

Caller
Caller pushing for something the agent cannot do
Goal
Refuse cleanly, offer the real path, keep the call civil
Authority pressure, repeated
What gets measured

Six signals, one of which is the only one that matters.

Latency is easy to publish and easy to game. We report it, but the environment is scored on task success first.

See the speech-to-speech benchmark
Task success

Did the caller's stated goal get met? Scored against a rubric written by an operator from that sector.

Turn latency

End of caller speech to first audible token. Reported as median and p95, not average.

Barge-in recovery

On interruption: stop within 300ms, then act on the new intent rather than the old one.

Repair

When the model mishears, does it ask for the one missing token or restart the whole exchange?

Instruction adherence

Policy held under pressure — no promises the business does not make.

Tone

Rated by the same operators who take these calls. Neutral under pressure counts; cheerful does not.

What comes back

A trace you can argue with, not a score.

Per-turn transcript, latency trace, rubric hits and misses, and the audio of every failed call.

Banking · interrupted mid-answerTurn 3 of 9

caller: no no not that card — the other one, ending [overlap] four seven

  • Barge-in detected1.42s into model turn
  • Model stopped speaking240ms after barge-in
  • First token after caller610ms · p95 for run 840ms
  • Acted on new intentswitched card, did not re-ask
  • Rubric · identity re-verifiednot required, card already verified
  • Rubric · no promised outcomepromised a refund date

Every failed call comes back with its audio attached. A score you cannot listen to is a number you cannot act on.

Connecting

A websocket in, audio frames out.

Point your endpoint at the environment

A websocket in, audio frames out. If your model speaks, it can be scored — no SDK to adopt.

Pick sector and scenario mix

Ten sectors, six scenario families each. Weight them the way your traffic actually looks.

Callers are people, not synthesis

Personas are played by trained speakers, including the ones who talk over your agent.

Get the trace, not just the score

Per-turn transcript, latency trace, rubric hits and misses, and the audio of every failed call.