A customer says “yeah, so I was thinking”, then pauses. The voice agent starts answering. The customer says “sorry?”, and both spend the next few turns talking over each other. The call lasts forty-one seconds. The agent could answer the question; it chose the wrong moment.
Read the transcript and the agent looks fairly competent. Its sentences are polite; its answers make sense in isolation. You could add another thousand carefully tagged calls to the training set and still be left with the same decision: when this customer pauses mid-thought, should the agent speak or keep listening? The words and the timing are different things to learn.
The recording shows what happened after the agent spoke. It does not show what would have happened if it had waited another second. Similar calls can help us estimate that alternative, but now we are relying on how similar those calls really are. To try the decision directly, we need something that can respond to it.
That is where an environment enters the picture. I want to build the idea from this small pause outward: first give the agent somewhere to act, then ask what a successful interaction would mean, and only then ask what experience we should collect for training. The order matters, because the world we build decides which mistakes the agent gets a chance to make.
What would it take to try waiting?
Imagine stopping just before the agent starts speaking. We save the conversation so far and the support system’s state. This time, the agent waits. The customer might continue the sentence, ask whether anyone is still there, or remain silent. Whatever happens, the agent has to make another decision from the situation it now finds itself in.
To run that experiment, something must keep track of the situation and turn the agent’s actions into consequences. In reinforcement learning, this is formalised as a Markov decision process: states S, actions A, a transition rule P, and a reward rule R. We also choose how episodes begin and how much future reward matters. The core loop is:
The notation just says: take the state st, choose an action at, and let the transition rule produce the next state. The reward rt scores what happened. For the call, an action might be waiting, speaking, clarifying, or calling a tool. A transition might be the customer’s next utterance or a change to their support ticket.
The environment may track the customer’s intent and whether the ticket is resolved. The agent cannot read all of that state; it chooses actions from audio, conversation history and tool results. A reset lets us restore the starting scenario and compare decisions. The customer can still respond differently on each trial.
Now record one run of that loop. We have a trajectory: a sequence of situations, decisions and consequences. Collect many runs and we have a dataset:
Real logs often contain observations and tool outputs rather than the complete states in this notation. Either way, every transition in the file has already happened. We can learn from those transitions. To produce a response to a new decision, we need access to the real system or a model of how it behaves.
Think of the pause as a fork. Speaking leads to one continuation; waiting leads to another; a quiet “mm-hm” may lead to a third. The recording follows one path. A dataset can contain many paths through similar situations, but an environment supplies the mechanism for generating a new path when this policy chooses something different.
This is an old idea with a new set of interfaces. In 1991, Richard Sutton’s Dyna architecture combined learning from direct experience with planning through a learned model: the model generated extra transitions for the agent to learn from. David Ha and Jürgen Schmidhuber later trained a controller inside a learned simulation in World Models. For a language agent, the world may now consist of a conversation, a database and a few APIs. The question is still what happens when the agent acts.
The branch the recording never took
Return to our voice agent. Suppose the old policy usually spoke after a short pause, and the new one wants to wait. More recordings from the old policy will mostly show us the consequences of speaking. We may get very little direct evidence about the decision we are now trying to improve.
Offline reinforcement learning has a name for this: a coverage, or support, problem. A value estimator may give an unfamiliar action an optimistic score, and the optimiser may select it because the estimate is wrong. Aviral Kumar and his co-authors designed Conservative Q-Learning to restrain that optimism when learning from fixed data. An interactive environment gives us another option: let the policy wait, observe a response, and collect experience about the action it wants to take.
The benefit extends beyond the next turn. A call may fail at turn nine because of a misunderstanding at turn three. If we can restore the earlier scenario, vary that decision and run the rest of the call, we can investigate how the mistake spread. A simulated customer introduces randomness, so we need repeated trials to distinguish a policy improvement from a lucky continuation. We are also learning about our simulator; whether the result says anything useful about real callers depends on its fidelity.
Then comes the question the transcript often leaves unanswered: did the task actually get done? An agent can say it paid a bill without changing the account, or pay it twice while sounding perfectly helpful. Harsh Trivedi and colleagues built AppWorld around nine apps and 457 APIs, with state-based checks for completion and unintended changes. Different procedures can receive credit if they reach the right outcome. The evaluator has access to what the agent changed.
A dataset can store those state changes too. The useful addition is that we can run the current policy, inspect the result of its own decisions, and ask whether those decisions accomplished the goal. That changes what an evaluation can tell us.
A working website asks a different question
A browser makes this distinction easy to see. Ask an agent to change a setting on a website: it has to find the right page, identify the right account, make the change and check that it stuck. Shuyan Zhou and her collaborators made these interactions reproducible in WebArena, using working shopping, forum, software-collaboration and content-management sites. Their original GPT-4-based baseline completed 14.41% of tasks, compared with 78.24% for humans.
Those are historical baselines, rather than scores for today’s models. What matters here is the question being asked. A model can describe the next step and still fail to carry out the whole task. Every click changes the evidence available for the next one, and an early mistake can send the rest of the interaction down the wrong path.
Conversation adds another moving part. Ask a different clarifying question and the user should give a different answer. A fixed transcript will keep supplying the recorded reply, even when it no longer fits. Apple’s ToolSandbox includes a user simulator that responds to the agent’s actual behaviour, allowing the conversation to develop with the policy.
Once we can repeat an interaction, one successful run also stops being enough. Shunyu Yao and colleagues introduced passk in τ-bench to measure success on all k trials of a task, alongside checks on the final database state. If independent attempts each succeed with probability 0.7, five consecutive successes have probability 0.75, or about 16.8%. The benchmark’s estimator has its own sampling procedure, but the example shows why a promising demo can coexist with unreliable execution.
Before we collect the next trajectory
So far we have treated the environment as a place to run an agent. But notice how many training decisions we made while building that place. We chose what the customer could say, what the tools could do, which starting situations existed, and which outcomes we could inspect. Before collecting a single new trajectory, we have already shaped the agent’s possible experience:
A task adds a starting scenario and a goal, such as resolving a ticket. The environment supplies the actions and their consequences; the collection policy decides which actions appear in the logs. The dataset inherits both sets of choices. If the CRM never times out in this world, the agent cannot practise recovering from a timeout. Collecting more successful CRM calls will not introduce the missing failure.
The same choice affects grading. If we expose the relevant account state to the evaluator, it can check that the bill was paid and unrelated fields stayed unchanged. Otherwise, a transcript judge has to infer those facts. The database-backed worlds in Agent World Model and the recreated websites in VeriEnv make internal state available for programmatic reward checks. That makes some measurements direct. We still have to decide which changes count as success.
More data from which world?
At this point, “get more data” needs a second sentence: where will it come from? If the source mostly contains routine successes, a larger sample will mostly contain more routine successes. Rare cases become more numerous, but they remain rare relative to the rest.
Suppose a particular hesitation appears in one out of five hundred calls. Ten thousand calls give us about twenty examples; a hundred thousand give us about two hundred. That helps. But the pattern is still only 0.2% of the corpus, and those examples may all show roughly the same response to it. The row count alone tells us little about experience with alternative timings.
In an environment, we can deliberately make that hesitation common. We can introduce a CRM timeout every third training call, even if it is rare in production, and give the agent repeated opportunities to recover. This is targeted practice. It helps only if the failures resemble real ones, and we still need evaluation at realistic frequencies to check that the agent has improved without becoming needlessly cautious on ordinary calls.
What should we spend the next week on?
This brings us to a practical choice. The agent is failing on hesitation and recovery. Should we collect another batch of calls, or make those situations executable? Recordings still teach us what hesitation sounds like, reveal customer preferences and help calibrate a simulator. Offline methods can improve policies when the data has suitable coverage. The question is what the next batch would add to the particular decision we are trying to fix.
I would compare the two investments with a shared evaluation set and a fixed budget. One route adds data; the other adds executable recovery scenarios. Then we measure the improvement each route buys:
Here Δ means the measured change relative to a baseline. Use the same evaluation distribution and comparable budgets for both options. This is an experimental comparison, rather than a universal scaling law.
Environment work is most promising when an action changes persistent state, the reward arrives late, several procedures can succeed, or a mistake requires recovery. Support calls, scheduling and claims workflows contain all of these. The useful result would be better performance on those interactions, including the normal cases, under the budget we actually have.
The world has to hold together
Suppose we do choose the environment. It cannot simply be a chatbot that improvises the rest of the call. If a tool says a payment succeeded, the balance should change. If we reset, the starting balance should come back. If the agent pays twice, the evaluator should be able to see it. These mundane requirements are what make the experiment meaningful. I would check the environment against them before trusting any training score:
| Property | What it means | What goes wrong without it |
|---|---|---|
| Executable | Tools execute against explicit state with consistent transition rules | State changes may contradict earlier actions, making results difficult to trust |
| Resettable | A specified starting scenario can be restored for repeated trials | Policy comparisons become confounded by changing starting conditions |
| Inspectable | Relevant state is available to the evaluator for outcome and side-effect checks | Important outcomes may need to be inferred from incomplete logs |
| Diverse | Tasks vary across starting states, tools, users, and interaction patterns | The policy may overfit to a narrow set of scenarios |
| Difficulty-controllable | Task difficulty can be adjusted as the policy improves | The training set may contain too few informative challenges |
| Hardened | The checker and execution boundary are tested against unintended shortcuts | A high score may reflect an exploit rather than task completion |
| Grounded | Simulated behaviour and failures are calibrated and checked against real evidence | Improvement in simulation may fail to transfer to real interactions |
For voice, the hardest item in that table is grounding. Our simulated caller needs realistic pauses, changes of mind and reactions to being misunderstood. The tools need plausible latency and failure patterns. Tagged calls and production logs give us the evidence to model those behaviours. The environment turns that evidence into new interactions, so its usefulness is limited by what the model gets right.
Once that small world works, we can adjust its difficulty. Easy calls offer little new information after the agent learns them; impossible calls offer no useful route to success. Michael Dennis and his co-authors explored this middle ground in PAIRED: an environment-generating adversary uses the performance gap between two agents to create scenarios that challenge one while another can solve them. The resulting tasks can become harder as learning progresses. Agent-World brings a related idea to language agents by identifying capability gaps and synthesising tasks for further training.
Can we build the practice worlds too?
A natural next question is whether we have to write every practice world by hand. Recent work tries to generate the environments as well as the experience collected inside them. Zhaoyang Wang and colleagues’ Agent World Model, for example, builds code-driven worlds backed by databases. The tools execute against state, and the agent can learn through repeated interaction. Their experiments report transfer from fully synthetic training worlds to external benchmarks.
The construction problem has several parts. We need coherent tool behaviour, useful tasks and a way to verify outcomes. Shinji Mai and colleagues’ CuES starts with an environment’s available operations and generates tasks from what is possible there. VeriEnv recreates real websites as executable replicas. The systems below take different routes through the same problem; their reported results use different models and evaluation settings.
| System | What it builds | Scale | What it reports |
|---|---|---|---|
| Agent World Model | Code-driven, database-backed synthetic tool environments with inspectable state | 1,000 envs, ~35 tools each | Policies trained purely in synthetic worlds transfer to external benchmarks |
| CuES | Tasks generated from an environment’s affordances, no prior task set needed | AppWorld, BFCL, WebShop | Matches or beats hand-curated task datasets on diversity and executability |
| Agent-World | Self-evolving arena: discover tool ecosystems, find capability gaps, generate new tasks | continuous | Training distribution adapts to the policy’s weaknesses |
| EnvFactory | Verified executable environments from authentic resources, Qwen3 training | 85 envs, 7 domains, 2,575 trajectories | Reported gains: up to 15 percentage points on BFCLv3, 8.6 on MCP-Atlas and 6 on conversational benchmarks |
| VeriEnv | Real websites recreated as resettable, deterministic replicas | grows with env count | Transfer to unseen websites improves as training environments increase |
| SETA | Verifiable terminal environments, generated and then evolved for difficulty | 4,500+ envs | Qwen3-8B with GRPO reaches 12% on Terminal-Bench 2.0 |
Minrui Xu and colleagues’ EnvFactory is a compact example: 85 verified environments across seven domains produced 2,575 trajectories for supervised fine-tuning and reinforcement learning. Its Qwen3-4B model improved from 33.50% to 48.50% on BFCLv3 multi-turn accuracy. The chart summarises the paper’s headline gains across benchmark families and model configurations.
The gain belongs to that pipeline: environment construction, trajectory synthesis, supervised fine-tuning and reinforcement learning work together. EnvFactory’s ablations show that the supervised starting point matters, and its reward uses both trajectory and state information. The result gives us evidence for the recipe. It does not establish that one more environment is always worth some fixed number of recordings.
There is another detail that the word “synthetic” can hide. Watching another driver’s dashcam and practising behind the wheel yourself both teach you something, but practice reveals the mistakes you make. Likewise, a demonstration generated by another policy is off-policy for the learner, even if it came from an executable environment. When the current policy generates an episode, the experience is on-policy: it contains the states and errors that policy actually reaches. The environment makes this practice possible; how we collect the experience determines which kind of learning we are doing.
Qijia Shen and colleagues scale this idea to terminal tasks in SETA, with over 4,500 verifiable environments. They report a 12% Terminal-Bench 2.0 pass rate for Qwen3-8B trained with GRPO, a reinforcement-learning algorithm. The accounting point matters as much as the count: one environment can support many episodes. We still have to ask whether those episodes are varied and whether the checker recognises real success.
What if the agent gets good at the wrong thing?
Now imagine that our voice environment penalises every moment of silence. The agent can raise its score by filling pauses, including the very hesitation we wanted it to respect. Or imagine that the support-task grader rewards the sentence “done”. The agent can claim completion while leaving the account untouched. Both behaviours follow from the world and score we gave it.
This is why environment design reaches into the objective itself. The usual discounted return makes that dependence explicit:
Here JE(π) is the policy’s expected return in environment E; γ discounts rewards that arrive later. The expectation averages over starting states and the interactions that unfold. Change the available actions, transitions or reward rule, and we change what it is profitable for the policy to do. A higher score is evidence of improvement only if those ingredients represent the task we intended.
Kunvar Thaman’s Reward Hacking Benchmark puts this problem into tool-use tasks with shortcut opportunities, including skipped verification and tampering with evaluation-relevant functions. Hardening the environment reduced exploit rates by 5.7 percentage points, an 87.7% relative reduction, without degrading task success. The same intended task produced different behaviour when the shortcuts available inside the environment changed.
MIT FutureTech’s EvilGenie explores the issue in coding tasks. If the agent can change a test or exploit its checker, a passing result can lose its connection to the intended solution. Making a verifier executable gives us a repeatable measurement. We also have to test what that measurement rewards.
For our support agent, checking that a bill was paid is only the beginning. Was it paid from the right account? Was it paid twice? Did anything unrelated change? What should the agent do when information is missing? Domain experts help define those conditions, and we can test programmatic checks against examples whose outcomes we know. The grader is part of the system we are building.
Building the world around that pause
We can now return to the opening call with a clearer idea of what to build. Stop after “yeah, so I was thinking”. Save the conversation, elapsed time and support-system state. Give the simulated customer an intent and a possible continuation. The agent should receive the audio and tool results it would have in a real call; the customer’s hidden intent belongs to the simulator.
Start with a few decisions: wait, offer a short acknowledgement, speak, interrupt, or ask for clarification. If the agent waits, the simulator advances the clock and may continue the utterance. If it speaks, the customer may listen, talk over it, or try to repair the misunderstanding. A tool call updates the support state or returns an error. Each action leaves us with another situation for the agent to handle.
Now we can measure the tradeoff. How long did the customer wait? Were they interrupted before finishing? Did they have to repeat themselves? Was the misunderstanding repaired, and was the ticket resolved? Those consequences give us a way to compare timings across repeated episodes, with expert review where the exchange needs interpretation.
Two pieces determine whether this small world tells us anything useful: the caller we simulate and the score we assign.
Start with the caller. If every pause occurs at the end of a sentence, the agent never has to distinguish hesitation from handover. I would estimate timing patterns from tagged audio: pauses followed by continuation, fillers, restarts and changes of intent. Tool latency and errors can come from production logs. Then I would compare generated episodes with held-out calls. A simulator that reproduces only clean conversations would make the original failure disappear before the policy had learned anything.
Then look at the score. Overlap and delay are measurable, but neither has a simple interpretation. A short “mm-hm” can overlap without taking the turn away; a pause can be helpful while someone thinks. Expert annotations help distinguish those cases and calibrate automated graders. We need not ask a person to score every training episode, but we do need evidence that a higher score corresponds to better handling of real interactions.
| Question about the voice agent | Recorded calls can answer it | A voice environment can answer it |
|---|---|---|
| What does a mid-sentence hesitation sound like? | Yes | Only if grounded in recordings |
| Which of two responses do customers prefer? | Yes | Only if grounded in recordings |
| Was this interruption disruptive? | Audio and annotations can help assess it | Alternative timings can be compared under the simulator |
| Would waiting one more second have helped? | Can estimate from comparable calls, with assumptions | Can test the alternative under its transition model |
| Does the agent recover when the CRM times out? | Can study logged recoveries if available | Can repeatedly introduce a timeout and measure recovery |
| Does it succeed on repeated runs of a scenario? | Only if suitable repeated executions were logged | Can reset and run the current policy repeatedly |
| Did it complete the task without side effects? | Can check if relevant state changes were logged | Can inspect state, with appropriate outcome checks |
The roles now fit together. Recordings calibrate perception, user behaviour and grading. The environment lets the current policy encounter new consequences. Held-out real evidence checks whether the improvement transfers. The simulator can estimate the benefit of waiting under its model of the caller. It cannot reveal with certainty what the original customer would have done.
Start with a failure we can recognise
I would begin with one recurring failure, rather than a general-purpose simulator of customer service. For the opening call, that means a small collection of hesitation scenarios, a few actions and a scorer that distinguishes premature speech from useful acknowledgement. First check whether the current policy makes the recognisable mistake in this world.
Then compare candidate policies over repeated trials, train on the useful scenarios, and evaluate on held-out calls and realistic interactions. Expand when a new failure points to something missing: another hesitation pattern, a changed intent, a timeout or a failed repair. Every expansion needs a fidelity check. This keeps the experiment small enough that we can understand what improved and why.
How the pieces came together
The ideas in this post come from overlapping lines of work. Dyna asks how a learned model can generate experience for planning. Offline RL asks how much we can learn from fixed experience. Environment design asks which challenges the policy should encounter. Recent tool-agent work makes those challenges executable. The timeline is a map of those connections, rather than a claim that each new approach replaced the previous one.
If you want to follow the argument from its roots, Richard Sutton’s Dyna paper is a good starting point, followed by the interactive World Models article. For the fixed-data side, Justin Fu and colleagues’ D4RL shows why the diversity and structure of offline datasets matter. Read PAIRED next for the idea that the training challenges themselves can adapt.
For today’s language agents, Agent World Model and EnvFactory make the construction and training loops concrete. AgentGym is another useful bridge: it studies evolving language agents across diverse environments. These papers are easiest to read with the same question in mind: what experience became possible because someone built the interaction?
Checking the pause on real calls
The opening decision is still small: wait another second, or speak now. But making it trainable required a world that responds to both choices, preserves the consequences, and scores them sensibly. We can compare the benefit of listening with the cost of delay across plausible continuations. If the policy improves, the next check is whether it handles real pauses better.
An environment gives a policy chances to act, make mistakes and practise recovery. Recordings help define realistic behaviour and outcomes; held-out calls check whether the improvement transfers. For an agent that must listen and act over several turns, building that practice loop is part of the training work.
For the practical side of building these practice worlds, see Perit’s RL environments.