{"id":896,"date":"2026-09-30T10:38:45","date_gmt":"2026-09-30T10:38:45","guid":{"rendered":"https:\/\/perit.ai\/blogs\/?p=896"},"modified":"2026-09-30T18:42:35","modified_gmt":"2026-09-30T18:42:35","slug":"rl-environments-voice-agents","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/rl-environments-voice-agents\/","title":{"rendered":"Why More Call Recordings Won&#8217;t Fix Your Voice Agent"},"content":{"rendered":"\n\n\r\n<div class=\"pp-post\" style=\"max-width: 50rem;margin: 0 auto;font-family: Verdana, Geneva, Tahoma, sans-serif;font-size: 1rem;line-height: 1.65;color: #0b0b0b\">\r\n<div class=\"meta\" style=\"font-size: 0.92rem;color: #52514e;margin-bottom: 1.6rem;padding-bottom: 0.65rem;border-bottom: 1px solid #e6e5e0\">One support-call pause, the branch a recording cannot show, and the world an agent needs to practise in<\/div>\r\n<p class=\"lede\" style=\"margin: 0 0 1rem;font-size: 1.08rem\">A customer says \u201cyeah, so I was thinking\u201d, then pauses. The voice agent starts answering. The customer says \u201csorry?\u201d, and both spend the next few turns talking over each other. The call lasts forty-one seconds. The agent could answer the question; it chose the wrong moment.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Read the transcript and the agent looks fairly competent. Its sentences are polite; its answers make sense in isolation. You could add another thousand carefully tagged calls to the training set and still be left with the same decision: when this customer pauses mid-thought, should the agent speak or keep listening? The words and the timing are different things to learn.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The recording shows what happened after the agent spoke. It does not show what would have happened if it had waited another second. Similar calls can help us estimate that alternative, but now we are relying on how similar those calls really are. To try the decision directly, we need something that can respond to it.<\/p>\r\n<p style=\"margin: 0 0 1rem\">That is where an environment enters the picture. I want to build the idea from this small pause outward: first give the agent somewhere to act, then ask what a successful interaction would mean, and only then ask what experience we should collect for training. The order matters, because the world we build decides which mistakes the agent gets a chance to make.<\/p>\r\n\r\n<div class=\"quote-block\" style=\"margin: 1.5rem 0;padding: 0.1rem 0 0.1rem 1rem;border-left: 3px solid #0b0b0b;font-style: italic;color: #0b0b0b\">An environment gives us a way to change a decision, run the interaction, and measure what follows. The training opportunity begins with that ability to try another branch.<\/div>\r\n<h2 id=\"what-would-it-take-to-try-waiting\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">What would it take to try waiting?<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Imagine stopping just before the agent starts speaking. We save the conversation so far and the support system&#8217;s state. This time, the agent waits. The customer might continue the sentence, ask whether anyone is still there, or remain silent. Whatever happens, the agent has to make another decision from the situation it now finds itself in.<\/p>\r\n<p style=\"margin: 0 0 1rem\">To run that experiment, something must keep track of the situation and turn the agent&#8217;s actions into consequences. In reinforcement learning, this is formalised as a Markov decision process: states <em>S<\/em>, actions <em>A<\/em>, a transition rule <em>P<\/em>, and a reward rule <em>R<\/em>. We also choose how episodes begin and how much future reward matters. The core loop is:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;background-color: #fcfcfb;border: 1px solid #e6e5e0;border-left: 3px solid #2a78d6;border-radius: 6px\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &apos;Times New Roman&apos;, Georgia, serif;font-size: 1.15rem;margin: 0.4rem 0;color: #0b0b0b\">s<sub>t+1<\/sub> \u223c P(\u00b7 | s<sub>t<\/sub>, a<sub>t<\/sub>), \u00a0 r<sub>t<\/sub> = R(s<sub>t<\/sub>, a<sub>t<\/sub>, s<sub>t+1<\/sub>)<\/span><\/div>\r\n<p style=\"margin: 0 0 1rem\">The notation just says: take the state <em>s<sub>t<\/sub><\/em>, choose an action <em>a<sub>t<\/sub><\/em>, and let the transition rule produce the next state. The reward <em>r<sub>t<\/sub><\/em> scores what happened. For the call, an action might be waiting, speaking, clarifying, or calling a tool. A transition might be the customer&#8217;s next utterance or a change to their support ticket.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The environment may track the customer\u2019s intent and whether the ticket is resolved. The agent cannot read all of that state; it chooses actions from audio, conversation history and tool results. A reset lets us restore the starting scenario and compare decisions. The customer can still respond differently on each trial.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">Two boxes, agent on the left and environment on the right. An arrow labelled action goes from agent to environment. Arrows labelled observation and reward return from environment to agent. A reset arrow loops back into the environment. A note says a dataset stores recorded transitions, while the environment generates responses to new actions. The environment loop A dataset stores recorded transitions. The environment generates responses to new actions. Agent chooses the next action Environment state \u00b7 transition rule reward \u00b7 reset action observation reward reset and try again<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: .65rem;font-size: .9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 1.<\/strong> The basic interaction loop. A reset lets us run another episode from a specified starting scenario, with a different action or policy.<\/figcaption><\/figure>\r\n<p style=\"margin: 0 0 1rem\">Now record one run of that loop. We have a <em>trajectory<\/em>: a sequence of situations, decisions and consequences. Collect many runs and we have a dataset:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;background-color: #fcfcfb;border: 1px solid #e6e5e0;border-left: 3px solid #2a78d6;border-radius: 6px\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &apos;Times New Roman&apos;, Georgia, serif;font-size: 1.15rem;margin: 0.4rem 0;color: #0b0b0b\">\u03c4 = (s<sub>0<\/sub>, a<sub>0<\/sub>, r<sub>0<\/sub>, s<sub>1<\/sub>, \u2026, s<sub>T<\/sub>), \u00a0 D = {\u03c4<sub>1<\/sub>, \u2026, \u03c4<sub>N<\/sub>}<\/span><\/div>\r\n<p style=\"margin: 0 0 1rem\">Real logs often contain observations and tool outputs rather than the complete states in this notation. Either way, every transition in the file has already happened. We can learn from those transitions. To produce a response to a new decision, we need access to the real system or a model of how it behaves.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Think of the pause as a fork. Speaking leads to one continuation; waiting leads to another; a quiet \u201cmm-hm\u201d may lead to a third. The recording follows one path. A dataset can contain many paths through similar situations, but an environment supplies the mechanism for generating a new path when this policy chooses something different.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">Left: a branching tree of states drawn in gray with a single path highlighted in blue, labelled the dataset. Right: the same tree with every branch drawn in blue, labelled the environment. One recorded trajectory versus the tree of reachable ones The dataset The environment one path, recorded. Other branchesare not observed in this recording.new branches can be generated,subject to the transition rules.<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: .65rem;font-size: .9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 2.<\/strong> A simplified branching picture. The highlighted path represents a recording; the environment can generate other paths according to its transition rules.<\/figcaption><\/figure>\r\n<p style=\"margin: 0 0 1rem\">This is an old idea with a new set of interfaces. In 1991, Richard Sutton&#8217;s <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/dl.acm.org\/doi\/10.1145\/122344.122377\" target=\"_blank\" rel=\"noopener\">Dyna architecture<\/a> combined learning from direct experience with planning through a learned model: the model generated extra transitions for the agent to learn from. David Ha and J\u00fcrgen Schmidhuber later trained a controller inside a learned simulation in <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/worldmodels.github.io\/\" target=\"_blank\" rel=\"noopener\">World Models<\/a>. For a language agent, the world may now consist of a conversation, a database and a few APIs. The question is still what happens when the agent acts.<\/p>\r\n\r\n<h2 id=\"the-branch-the-recording-never-took\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">The branch the recording never took<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Return to our voice agent. Suppose the old policy usually spoke after a short pause, and the new one wants to wait. More recordings from the old policy will mostly show us the consequences of speaking. We may get very little direct evidence about the decision we are now trying to improve.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Offline reinforcement learning has a name for this: a coverage, or <em>support<\/em>, problem. A value estimator may give an unfamiliar action an optimistic score, and the optimiser may select it because the estimate is wrong. Aviral Kumar and his co-authors designed <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2006.04779\" target=\"_blank\" rel=\"noopener\">Conservative Q-Learning<\/a> to restrain that optimism when learning from fixed data. An interactive environment gives us another option: let the policy wait, observe a response, and collect experience about the action it wants to take.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The benefit extends beyond the next turn. A call may fail at turn nine because of a misunderstanding at turn three. If we can restore the earlier scenario, vary that decision and run the rest of the call, we can investigate how the mistake spread. A simulated customer introduces randomness, so we need repeated trials to distinguish a policy improvement from a lucky continuation. We are also learning about our simulator; whether the result says anything useful about real callers depends on its fidelity.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Then comes the question the transcript often leaves unanswered: did the task actually get done? An agent can say it paid a bill without changing the account, or pay it twice while sounding perfectly helpful. Harsh Trivedi and colleagues built <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2407.18901\" target=\"_blank\" rel=\"noopener\">AppWorld<\/a> around nine apps and 457 APIs, with state-based checks for completion and unintended changes. Different procedures can receive credit if they reach the right outcome. The evaluator has access to what the agent changed.<\/p>\r\n<p style=\"margin: 0 0 1rem\">A dataset can store those state changes too. The useful addition is that we can run the current policy, inspect the result of its own decisions, and ask whether those decisions accomplished the goal. That changes what an evaluation can tell us.<\/p>\r\n\r\n<h2 id=\"a-working-website-asks-a-different-question\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">A working website asks a different question<\/h2>\r\n<p style=\"margin: 0 0 1rem\">A browser makes this distinction easy to see. Ask an agent to change a setting on a website: it has to find the right page, identify the right account, make the change and check that it stuck. Shuyan Zhou and her collaborators made these interactions reproducible in <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2307.13854\" target=\"_blank\" rel=\"noopener\">WebArena<\/a>, using working shopping, forum, software-collaboration and content-management sites. Their original GPT-4-based baseline completed 14.41% of tasks, compared with 78.24% for humans.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">Two horizontal bars. Humans at 78.24 percent, GPT-4 baseline at 14.41 percent, on a common 0 to 100 axis. End-to-end success on WebArena (%) Original evaluation: humans versus the GPT-4-based baseline Participant 0255075100 Tasks completed end-to-end (%) Humans 78.24% GPT-4 baseline 14.41%<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: .65rem;font-size: .9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 3.<\/strong> In <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2307.13854\" target=\"_blank\" rel=\"noopener\">WebArena&#8217;s original evaluation<\/a>, the GPT-4-based agent completed 14.41% of tasks and humans completed 78.24%. The gap concerns full task execution, with all the intermediate decisions included.<\/figcaption><\/figure>\r\n<p style=\"margin: 0 0 1rem\">Those are historical baselines, rather than scores for today&#8217;s models. What matters here is the question being asked. A model can describe the next step and still fail to carry out the whole task. Every click changes the evidence available for the next one, and an early mistake can send the rest of the interaction down the wrong path.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Conversation adds another moving part. Ask a different clarifying question and the user should give a different answer. A fixed transcript will keep supplying the recorded reply, even when it no longer fits. Apple&#8217;s <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/machinelearning.apple.com\/research\/toolsandbox-stateful-conversational-llm-benchmark\" target=\"_blank\" rel=\"noopener\">ToolSandbox<\/a> includes a user simulator that responds to the agent&#8217;s actual behaviour, allowing the conversation to develop with the policy.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Once we can repeat an interaction, one successful run also stops being enough. Shunyu Yao and colleagues introduced <em>pass<sup>k<\/sup><\/em> in <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2406.12045\" target=\"_blank\" rel=\"noopener\">\u03c4-bench<\/a> to measure success on all <em>k<\/em> trials of a task, alongside checks on the final database state. If independent attempts each succeed with probability 0.7, five consecutive successes have probability 0.7<sup>5<\/sup>, or about 16.8%. The benchmark&#8217;s estimator has its own sampling procedure, but the example shows why a promising demo can coexist with unreliable execution.<\/p>\r\n\r\n<h2 id=\"before-we-collect-the-next-trajectory\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">Before we collect the next trajectory<\/h2>\r\n<p style=\"margin: 0 0 1rem\">So far we have treated the environment as a place to run an agent. But notice how many training decisions we made while building that place. We chose what the customer could say, what the tools could do, which starting situations existed, and which outcomes we could inspect. Before collecting a single new trajectory, we have already shaped the agent&#8217;s possible experience:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;background-color: #fcfcfb;border: 1px solid #e6e5e0;border-left: 3px solid #2a78d6;border-radius: 6px\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &apos;Times New Roman&apos;, Georgia, serif;font-size: 1.15rem;margin: 0.4rem 0;color: #0b0b0b\">environment \u2192 possible tasks \u2192 reachable trajectories \u2192 observable consequences \u2192 rewards \u2192 policy<\/span><\/div>\r\n<p style=\"margin: 0 0 1rem\">A task adds a starting scenario and a goal, such as resolving a ticket. The environment supplies the actions and their consequences; the collection policy decides which actions appear in the logs. The dataset inherits both sets of choices. If the CRM never times out in this world, the agent cannot practise recovering from a timeout. Collecting more successful CRM calls will not introduce the missing failure.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The same choice affects grading. If we expose the relevant account state to the evaluator, it can check that the bill was paid and unrelated fields stayed unchanged. Otherwise, a transcript judge has to infer those facts. The database-backed worlds in <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2602.10090\" target=\"_blank\" rel=\"noopener\">Agent World Model<\/a> and the recreated websites in <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2603.10505\" target=\"_blank\" rel=\"noopener\">VeriEnv<\/a> make internal state available for programmatic reward checks. That makes some measurements direct. We still have to decide which changes count as success.<\/p>\r\n\r\n<h2 id=\"more-data-from-which-world\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">More data from which world?<\/h2>\r\n<p style=\"margin: 0 0 1rem\">At this point, \u201cget more data\u201d needs a second sentence: where will it come from? If the source mostly contains routine successes, a larger sample will mostly contain more routine successes. Rare cases become more numerous, but they remain rare relative to the rest.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Suppose a particular hesitation appears in one out of five hundred calls. Ten thousand calls give us about twenty examples; a hundred thousand give us about two hundred. That helps. But the pattern is still only 0.2% of the corpus, and those examples may all show roughly the same response to it. The row count alone tells us little about experience with alternative timings.<\/p>\r\n<p style=\"margin: 0 0 1rem\">In an environment, we can deliberately make that hesitation common. We can introduce a CRM timeout every third training call, even if it is rare in production, and give the agent repeated opportunities to recover. This is targeted practice. It helps only if the failures resemble real ones, and we still need evaluation at realistic frequencies to check that the agent has improved without becoming needlessly cautious on ordinary calls.<\/p>\r\n\r\n<h2 id=\"what-should-we-spend-the-next-week-on\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">What should we spend the next week on?<\/h2>\r\n<p style=\"margin: 0 0 1rem\">This brings us to a practical choice. The agent is failing on hesitation and recovery. Should we collect another batch of calls, or make those situations executable? Recordings still teach us what hesitation sounds like, reveal customer preferences and help calibrate a simulator. Offline methods can improve policies when the data has suitable coverage. The question is what the next batch would add to the particular decision we are trying to fix.<\/p>\r\n<p style=\"margin: 0 0 1rem\">I would compare the two investments with a shared evaluation set and a fixed budget. One route adds data; the other adds executable recovery scenarios. Then we measure the improvement each route buys:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;background-color: #fcfcfb;border: 1px solid #e6e5e0;border-left: 3px solid #2a78d6;border-radius: 6px\">\r\n\r\n<span class=\"eq\" style=\"display: block;text-align: center;font-family: &apos;Times New Roman&apos;, Georgia, serif;font-size: 1.15rem;margin: 0.4rem 0;color: #0b0b0b\">\u0394 success \/ cost of environment work \u00a0 compared with \u00a0 \u0394 success \/ cost of additional data<\/span>\r\n<p style=\"margin: 0.6rem 0 0;color: #52514e;font-size: 0.92rem\">Here \u0394 means the measured change relative to a baseline. Use the same evaluation distribution and comparable budgets for both options. This is an experimental comparison, rather than a universal scaling law.<\/p>\r\n\r\n<\/div>\r\n<p style=\"margin: 0 0 1rem\">Environment work is most promising when an action changes persistent state, the reward arrives late, several procedures can succeed, or a mistake requires recovery. Support calls, scheduling and claims workflows contain all of these. The useful result would be better performance on those interactions, including the normal cases, under the budget we actually have.<\/p>\r\n\r\n<h2 id=\"the-world-has-to-hold-together\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">The world has to hold together<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Suppose we do choose the environment. It cannot simply be a chatbot that improvises the rest of the call. If a tool says a payment succeeded, the balance should change. If we reset, the starting balance should come back. If the agent pays twice, the evaluator should be able to see it. These mundane requirements are what make the experiment meaningful. I would check the environment against them before trusting any training score:<\/p>\r\n\r\n<div class=\"table-wrap\" style=\"max-width: 100%;overflow: auto;margin: 1.5rem 0\">\r\n<table style=\"width: 100%;min-width: 640px;border-collapse: collapse;font-size: 0.92rem;line-height: 1.5\">\r\n<thead>\r\n<tr>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">Property<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">What it means<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">What goes wrong without it<\/th>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Executable<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Tools execute against explicit state with consistent transition rules<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">State changes may contradict earlier actions, making results difficult to trust<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Resettable<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">A specified starting scenario can be restored for repeated trials<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Policy comparisons become confounded by changing starting conditions<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Inspectable<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Relevant state is available to the evaluator for outcome and side-effect checks<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Important outcomes may need to be inferred from incomplete logs<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Diverse<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Tasks vary across starting states, tools, users, and interaction patterns<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">The policy may overfit to a narrow set of scenarios<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Difficulty-controllable<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Task difficulty can be adjusted as the policy improves<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">The training set may contain too few informative challenges<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Hardened<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">The checker and execution boundary are tested against unintended shortcuts<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">A high score may reflect an exploit rather than task completion<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Grounded<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Simulated behaviour and failures are calibrated and checked against real evidence<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Improvement in simulation may fail to transfer to real interactions<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\n<p style=\"margin: 0 0 1rem\">For voice, the hardest item in that table is grounding. Our simulated caller needs realistic pauses, changes of mind and reactions to being misunderstood. The tools need plausible latency and failure patterns. Tagged calls and production logs give us the evidence to model those behaviours. The environment turns that evidence into new interactions, so its usefulness is limited by what the model gets right.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Once that small world works, we can adjust its difficulty. Easy calls offer little new information after the agent learns them; impossible calls offer no useful route to success. Michael Dennis and his co-authors explored this middle ground in <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2012.02096\" target=\"_blank\" rel=\"noopener\">PAIRED<\/a>: an environment-generating adversary uses the performance gap between two agents to create scenarios that challenge one while another can solve them. The resulting tasks can become harder as learning progresses. <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2604.18292\" target=\"_blank\" rel=\"noopener\">Agent-World<\/a> brings a related idea to language agents by identifying capability gaps and synthesising tasks for further training.<\/p>\r\n\r\n<h2 id=\"can-we-build-the-practice-worlds-too\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">Can we build the practice worlds too?<\/h2>\r\n<p style=\"margin: 0 0 1rem\">A natural next question is whether we have to write every practice world by hand. Recent work tries to generate the environments as well as the experience collected inside them. Zhaoyang Wang and colleagues&#8217; <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2602.10090\" target=\"_blank\" rel=\"noopener\">Agent World Model<\/a>, for example, builds code-driven worlds backed by databases. The tools execute against state, and the agent can learn through repeated interaction. Their experiments report transfer from fully synthetic training worlds to external benchmarks.<\/p>\r\n<p style=\"margin: 0 0 1rem\">The construction problem has several parts. We need coherent tool behaviour, useful tasks and a way to verify outcomes. Shinji Mai and colleagues&#8217; <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2512.01311\" target=\"_blank\" rel=\"noopener\">CuES<\/a> starts with an environment&#8217;s available operations and generates tasks from what is possible there. <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2603.10505\" target=\"_blank\" rel=\"noopener\">VeriEnv<\/a> recreates real websites as executable replicas. The systems below take different routes through the same problem; their reported results use different models and evaluation settings.<\/p>\r\n\r\n<div class=\"table-wrap\" style=\"max-width: 100%;overflow: auto;margin: 1.5rem 0\">\r\n<table style=\"width: 100%;min-width: 640px;border-collapse: collapse;font-size: 0.92rem;line-height: 1.5\">\r\n<thead>\r\n<tr>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">System<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">What it builds<\/th>\r\n<th class=\"num\" style=\"vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b;text-align: right\">Scale<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">What it reports<\/th>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\"><a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2602.10090\" target=\"_blank\" rel=\"noopener\">Agent World Model<\/a><\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Code-driven, database-backed synthetic tool environments with inspectable state<\/td>\r\n<td class=\"num\" style=\"vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;text-align: right;white-space: nowrap\">1,000 envs, ~35 tools each<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Policies trained purely in synthetic worlds transfer to external benchmarks<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\"><a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2512.01311\" target=\"_blank\" rel=\"noopener\">CuES<\/a><\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Tasks generated from an environment&#8217;s affordances, no prior task set needed<\/td>\r\n<td class=\"num\" style=\"vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;text-align: right;white-space: nowrap\">AppWorld, BFCL, WebShop<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Matches or beats hand-curated task datasets on diversity and executability<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\"><a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2604.18292\" target=\"_blank\" rel=\"noopener\">Agent-World<\/a><\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Self-evolving arena: discover tool ecosystems, find capability gaps, generate new tasks<\/td>\r\n<td class=\"num\" style=\"vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;text-align: right;white-space: nowrap\">continuous<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Training distribution adapts to the policy&#8217;s weaknesses<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\"><a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2605.18703\" target=\"_blank\" rel=\"noopener\">EnvFactory<\/a><\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Verified executable environments from authentic resources, Qwen3 training<\/td>\r\n<td class=\"num\" style=\"vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;text-align: right;white-space: nowrap\">85 envs, 7 domains, 2,575 trajectories<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Reported gains: up to 15 percentage points on BFCLv3, 8.6 on MCP-Atlas and 6 on conversational benchmarks<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\"><a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2603.10505\" target=\"_blank\" rel=\"noopener\">VeriEnv<\/a><\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Real websites recreated as resettable, deterministic replicas<\/td>\r\n<td class=\"num\" style=\"vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;text-align: right;white-space: nowrap\">grows with env count<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Transfer to unseen websites improves as training environments increase<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\"><a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2607.10891\" target=\"_blank\" rel=\"noopener\">SETA<\/a><\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Verifiable terminal environments, generated and then evolved for difficulty<\/td>\r\n<td class=\"num\" style=\"vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b;text-align: right;white-space: nowrap\">4,500+ envs<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Qwen3-8B with GRPO reaches 12% on Terminal-Bench 2.0<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\n<p style=\"margin: 0 0 1rem\">Minrui Xu and colleagues&#8217; <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2605.18703\" target=\"_blank\" rel=\"noopener\">EnvFactory<\/a> is a compact example: 85 verified environments across seven domains produced 2,575 trajectories for supervised fine-tuning and reinforcement learning. Its Qwen3-4B model improved from 33.50% to 48.50% on BFCLv3 multi-turn accuracy. The chart summarises the paper&#8217;s headline gains across benchmark families and model configurations.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">Three horizontal bars showing percentage-point improvements: BFCLv3 up to 15, MCP-Atlas 8.6, conversational benchmarks 6. EnvFactory: improvement on benchmark (percentage points) Headline gains from 2,575 environment-grounded trajectories Benchmark 05101520 Improvement (percentage points) BFCLv3 up to +15 MCP-Atlas +8.6 Conversational +6<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: .65rem;font-size: .9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 4.<\/strong> The <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2605.18703\" target=\"_blank\" rel=\"noopener\">EnvFactory authors<\/a> report these headline gains across benchmark families and model configurations. The bars summarise that pipeline&#8217;s results; they do not isolate the effect of environment quality from the rest of training.<\/figcaption><\/figure>\r\n<p style=\"margin: 0 0 1rem\">The gain belongs to that pipeline: environment construction, trajectory synthesis, supervised fine-tuning and reinforcement learning work together. EnvFactory&#8217;s ablations show that the supervised starting point matters, and its reward uses both trajectory and state information. The result gives us evidence for the recipe. It does not establish that one more environment is always worth some fixed number of recordings.<\/p>\r\n<p style=\"margin: 0 0 1rem\">There is another detail that the word \u201csynthetic\u201d can hide. Watching another driver&#8217;s dashcam and practising behind the wheel yourself both teach you something, but practice reveals the mistakes <em>you<\/em> make. Likewise, a demonstration generated by another policy is off-policy for the learner, even if it came from an executable environment. When the current policy generates an episode, the experience is <em>on-policy<\/em>: it contains the states and errors that policy actually reaches. The environment makes this practice possible; how we collect the experience determines which kind of learning we are doing.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Qijia Shen and colleagues scale this idea to terminal tasks in <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2607.10891\" target=\"_blank\" rel=\"noopener\">SETA<\/a>, with over 4,500 verifiable environments. They report a 12% Terminal-Bench 2.0 pass rate for Qwen3-8B trained with GRPO, a reinforcement-learning algorithm. The accounting point matters as much as the count: one environment can support many episodes. We still have to ask whether those episodes are varied and whether the checker recognises real success.<\/p>\r\n\r\n<h2 id=\"what-if-the-agent-gets-good-at-the-wrong-thing\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">What if the agent gets good at the wrong thing?<\/h2>\r\n<p style=\"margin: 0 0 1rem\">Now imagine that our voice environment penalises every moment of silence. The agent can raise its score by filling pauses, including the very hesitation we wanted it to respect. Or imagine that the support-task grader rewards the sentence \u201cdone\u201d. The agent can claim completion while leaving the account untouched. Both behaviours follow from the world and score we gave it.<\/p>\r\n<p style=\"margin: 0 0 1rem\">This is why environment design reaches into the objective itself. The usual discounted return makes that dependence explicit:<\/p>\r\n\r\n<div class=\"equation-note\" style=\"margin: 1.5rem 0;padding: 0.9rem 1rem;background-color: #fcfcfb;border: 1px solid #e6e5e0;border-left: 3px solid #2a78d6;border-radius: 6px\"><span class=\"eq\" style=\"display: block;text-align: center;font-family: &apos;Times New Roman&apos;, Georgia, serif;font-size: 1.15rem;margin: 0.4rem 0;color: #0b0b0b\">J<sub>E<\/sub>(\u03c0) = \ud835\udd3c<sub>\u03c4 \u223c (\u03c0, E)<\/sub>[\u03a3<sub>t=0<\/sub><sup>T\u22121<\/sup> \u03b3<sup>t<\/sup> r<sub>t<\/sub>]<\/span><\/div>\r\n<p style=\"margin: 0 0 1rem\">Here <em>J<sub>E<\/sub>(\u03c0)<\/em> is the policy&#8217;s expected return in environment <em>E<\/em>; <em>\u03b3<\/em> discounts rewards that arrive later. The expectation averages over starting states and the interactions that unfold. Change the available actions, transitions or reward rule, and we change what it is profitable for the policy to do. A higher score is evidence of improvement only if those ingredients represent the task we intended.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Kunvar Thaman&#8217;s <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2605.02964\" target=\"_blank\" rel=\"noopener\">Reward Hacking Benchmark<\/a> puts this problem into tool-use tasks with shortcut opportunities, including skipped verification and tampering with evaluation-relevant functions. Hardening the environment reduced exploit rates by 5.7 percentage points, an 87.7% relative reduction, without degrading task success. The same intended task produced different behaviour when the shortcuts available inside the environment changed.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">Two horizontal bars: exploit rate before hardening about 6.5 percent, after hardening about 0.8 percent, a reduction of 5.7 percentage points or 87.7 percent relative. Exploit rate on the Reward Hacking Benchmark (%) Approximate endpoints, reconstructed from the reported reductions Environment 02468 Exploit rate (%) Original environment \u22486.5% Hardened environment \u22480.8% \u22125.7 pts (\u221287.7% relative)<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: .65rem;font-size: .9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 5.<\/strong> In <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2605.02964\" target=\"_blank\" rel=\"noopener\">Thaman&#8217;s hardening experiment<\/a>, exploit rates fell by 5.7 percentage points, or 87.7% relative. The approximate endpoints shown here are reconstructed from those two reported reductions.<\/figcaption><\/figure>\r\n<p style=\"margin: 0 0 1rem\">MIT FutureTech&#8217;s <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/futuretech.mit.edu\/publication\/evilgenie-a-reward-hacking-benchmark\" target=\"_blank\" rel=\"noopener\">EvilGenie<\/a> explores the issue in coding tasks. If the agent can change a test or exploit its checker, a passing result can lose its connection to the intended solution. Making a verifier executable gives us a repeatable measurement. We also have to test what that measurement rewards.<\/p>\r\n<p style=\"margin: 0 0 1rem\">For our support agent, checking that a bill was paid is only the beginning. Was it paid from the right account? Was it paid twice? Did anything unrelated change? What should the agent do when information is missing? Domain experts help define those conditions, and we can test programmatic checks against examples whose outcomes we know. The grader is part of the system we are building.<\/p>\r\n\r\n<h2 id=\"building-the-world-around-that-pause\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">Building the world around that pause<\/h2>\r\n<p style=\"margin: 0 0 1rem\">We can now return to the opening call with a clearer idea of what to build. Stop after \u201cyeah, so I was thinking\u201d. Save the conversation, elapsed time and support-system state. Give the simulated customer an intent and a possible continuation. The agent should receive the audio and tool results it would have in a real call; the customer&#8217;s hidden intent belongs to the simulator.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Start with a few decisions: wait, offer a short acknowledgement, speak, interrupt, or ask for clarification. If the agent waits, the simulator advances the clock and may continue the utterance. If it speaks, the customer may listen, talk over it, or try to repair the misunderstanding. A tool call updates the support state or returns an error. Each action leaves us with another situation for the agent to handle.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Now we can measure the tradeoff. How long did the customer wait? Were they interrupted before finishing? Did they have to repeat themselves? Was the misunderstanding repaired, and was the ticket resolved? Those consequences give us a way to compare timings across repeated episodes, with expert review where the exchange needs interpretation.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">On the left, a box representing the call state. Five arrows fan out to five action boxes: wait, backchannel, speak, interrupt, clarify. From those, arrows converge on a box listing measured consequences: latency, false interruption, repetition, repair, task completion, tool failure and recovery, expert grade. A voice environment: one call state, five actions, measured consequences Call state audio \u00b7 transcript intent \u00b7 tool results time since last word wait backchannel (\u201cmm-hm\u201d) speak interrupt clarifyWhat it costlatency felt by the customerfalse interruptioncustomer had to repeatmisunderstanding repairedtask completedtool failed and recoveredexpert gradeone scenario, five possible actionsmeasure what follows each action<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: .65rem;font-size: .9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 6.<\/strong> Possible actions from one simulated call scenario, and the measurements used to compare their consequences. Repeated trials are needed when user behaviour is stochastic.<\/figcaption><\/figure>\r\n<p style=\"margin: 0 0 1rem\">Two pieces determine whether this small world tells us anything useful: the caller we simulate and the score we assign.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Start with the caller. If every pause occurs at the end of a sentence, the agent never has to distinguish hesitation from handover. I would estimate timing patterns from tagged audio: pauses followed by continuation, fillers, restarts and changes of intent. Tool latency and errors can come from production logs. Then I would compare generated episodes with held-out calls. A simulator that reproduces only clean conversations would make the original failure disappear before the policy had learned anything.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Then look at the score. Overlap and delay are measurable, but neither has a simple interpretation. A short \u201cmm-hm\u201d can overlap without taking the turn away; a pause can be helpful while someone thinks. Expert annotations help distinguish those cases and calibrate automated graders. We need not ask a person to score every training episode, but we do need evidence that a higher score corresponds to better handling of real interactions.<\/p>\r\n\r\n<div class=\"table-wrap\" style=\"max-width: 100%;overflow: auto;margin: 1.5rem 0\">\r\n<table style=\"width: 100%;min-width: 640px;border-collapse: collapse;font-size: 0.92rem;line-height: 1.5\">\r\n<thead>\r\n<tr>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">Question about the voice agent<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">Recorded calls can answer it<\/th>\r\n<th style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 2px solid #0b0b0b;font-weight: bold;color: #0b0b0b\">A voice environment can answer it<\/th>\r\n<\/tr>\r\n<\/thead>\r\n<tbody>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">What does a mid-sentence hesitation sound like?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Yes<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Only if grounded in recordings<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Which of two responses do customers prefer?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Yes<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Only if grounded in recordings<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Was this interruption disruptive?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Audio and annotations can help assess it<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Alternative timings can be compared under the simulator<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Would waiting one more second have helped?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Can estimate from comparable calls, with assumptions<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Can test the alternative under its transition model<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Does the agent recover when the CRM times out?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Can study logged recoveries if available<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Can repeatedly introduce a timeout and measure recovery<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Does it succeed on repeated runs of a scenario?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Only if suitable repeated executions were logged<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Can reset and run the current policy repeatedly<\/td>\r\n<\/tr>\r\n<tr>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Did it complete the task without side effects?<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Can check if relevant state changes were logged<\/td>\r\n<td style=\"text-align: left;vertical-align: top;padding: 0.55rem 0.6rem;border-bottom: 1px solid #e6e5e0;color: #0b0b0b\">Can inspect state, with appropriate outcome checks<\/td>\r\n<\/tr>\r\n<\/tbody>\r\n<\/table>\r\n<\/div>\r\n<p style=\"margin: 0 0 1rem\">The roles now fit together. Recordings calibrate perception, user behaviour and grading. The environment lets the current policy encounter new consequences. Held-out real evidence checks whether the improvement transfers. The simulator can estimate the benefit of waiting under its model of the caller. It cannot reveal with certainty what the original customer would have done.<\/p>\r\n\r\n<h2 id=\"start-with-a-failure-we-can-recognise\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">Start with a failure we can recognise<\/h2>\r\n<p style=\"margin: 0 0 1rem\">I would begin with one recurring failure, rather than a general-purpose simulator of customer service. For the opening call, that means a small collection of hesitation scenarios, a few actions and a scorer that distinguishes premature speech from useful acknowledgement. First check whether the current policy makes the recognisable mistake in this world.<\/p>\r\n<p style=\"margin: 0 0 1rem\">Then compare candidate policies over repeated trials, train on the useful scenarios, and evaluate on held-out calls and realistic interactions. Expand when a new failure points to something missing: another hesitation pattern, a changed intent, a timeout or a failed repair. Every expansion needs a fidelity check. This keeps the experiment small enough that we can understand what improved and why.<\/p>\r\n\r\n<h2 id=\"how-the-pieces-came-together\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">How the pieces came together<\/h2>\r\n<p style=\"margin: 0 0 1rem\">The ideas in this post come from overlapping lines of work. Dyna asks how a learned model can generate experience for planning. Offline RL asks how much we can learn from fixed experience. Environment design asks which challenges the policy should encounter. Recent tool-agent work makes those challenges executable. The timeline is a map of those connections, rather than a claim that each new approach replaced the previous one.<\/p>\r\n\r\n<figure class=\"viz\" style=\"margin: 2.2rem 0\">\r\n<div class=\"viz-card\" style=\"border: 1px solid #e6e5e0;border-radius: 10px;background-color: #fcfcfb;padding: 14px 14px 8px;overflow: auto\">A timeline of selected papers from overlapping strands: offline datasets, designed task distributions, executable worlds, and grounded training. Key papers are marked as dots; positions are not to scale. Selected papers, 1991\u20132026 Selected contributions from overlapping research strands. Positions are not to scale. Dyna 1991 World Models 2018 D4RL \u00b7 CQL \u00b7 PAIRED 2020 WebArena 2023 AppWorld \u00b7 \u03c4-bench 2024 CuES \u00b7 AgentGym 2025 AWM \u00b7 EnvFactory \u00b7 SETA 2026offlinedatasetsdesigned taskdistributionsexecutableworldsgroundedtraining<\/div>\r\n<figcaption class=\"viz-caption\" style=\"margin-top: .65rem;font-size: .9rem;line-height: 1.55;color: #52514e\"><strong style=\"color: #0b0b0b\">Figure 7.<\/strong> Selected contributions to learning from fixed data, designing training distributions, and building executable agent environments. The strands overlap, and the positions are not to scale.<\/figcaption><\/figure>\r\n<p style=\"margin: 0 0 1rem\">If you want to follow the argument from its roots, Richard Sutton&#8217;s <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/dl.acm.org\/doi\/10.1145\/122344.122377\" target=\"_blank\" rel=\"noopener\">Dyna paper<\/a> is a good starting point, followed by the interactive <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/worldmodels.github.io\/\" target=\"_blank\" rel=\"noopener\">World Models article<\/a>. For the fixed-data side, Justin Fu and colleagues&#8217; <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2004.07219\" target=\"_blank\" rel=\"noopener\">D4RL<\/a> shows why the diversity and structure of offline datasets matter. Read <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2012.02096\" target=\"_blank\" rel=\"noopener\">PAIRED<\/a> next for the idea that the training challenges themselves can adapt.<\/p>\r\n<p style=\"margin: 0 0 1rem\">For today&#8217;s language agents, <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2602.10090\" target=\"_blank\" rel=\"noopener\">Agent World Model<\/a> and <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/arxiv.org\/abs\/2605.18703\" target=\"_blank\" rel=\"noopener\">EnvFactory<\/a> make the construction and training loops concrete. <a style=\"color: #2a78d6;text-decoration: underline\" href=\"https:\/\/aclanthology.org\/2025.acl-long.1355\/\" target=\"_blank\" rel=\"noopener\">AgentGym<\/a> is another useful bridge: it studies evolving language agents across diverse environments. These papers are easiest to read with the same question in mind: what experience became possible because someone built the interaction?<\/p>\r\n\r\n<h2 id=\"back-to-the-pause\" style=\"font-size: 1.25rem;line-height: 1.35;margin: 2.4rem 0 0.65rem;font-weight: bold;color: #0b0b0b\">Checking the pause on real calls<\/h2>\r\n<p style=\"margin: 0 0 1rem\">The opening decision is still small: wait another second, or speak now. But making it trainable required a world that responds to both choices, preserves the consequences, and scores them sensibly. We can compare the benefit of listening with the cost of delay across plausible continuations. If the policy improves, the next check is whether it handles real pauses better.<\/p>\r\n<p style=\"margin: 0 0 1rem\">An environment gives a policy chances to act, make mistakes and practise recovery. Recordings help define realistic behaviour and outcomes; held-out calls check whether the improvement transfers. For an agent that must listen and act over several turns, building that practice loop is part of the training work.<\/p>\r\n\r\n<p style=\"margin: 1.4rem 0;font-size: 1rem;line-height: 1.68\">For the practical side of building these practice worlds, see <a href=\"https:\/\/perit.ai\/rl-environments\">Perit\u2019s RL environments<\/a>.<\/p>\n<footer class=\"post-note\" style=\"margin-top: 3rem;padding-top: 1rem;border-top: 1px solid #e6e5e0;font-size: 0.88rem;line-height: 1.5;color: #52514e\" aria-label=\"Notes on the examples and figures\">\r\n<p style=\"margin: 0 0 1rem\">The opening call is lightly anonymised. The voice environment described here is a proposed experiment, rather than a result from a deployed system. Benchmark figures refer to the linked authors&#8217; evaluation settings. The reward-hacking chart uses reconstructed approximate endpoints; the diagrams are illustrative, and the timeline is not to scale.<\/p>\r\n\r\n<\/footer><\/div>\r\n\n\n","protected":false},"excerpt":{"rendered":"<p>A recorded call shows one outcome. An RL environment lets a voice agent try other decisions, with real evidence to check whether practice transfers.<\/p>\n","protected":false},"author":2,"featured_media":897,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[22],"tags":[],"class_list":["post-896","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/896","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=896"}],"version-history":[{"count":5,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/896\/revisions"}],"predecessor-version":[{"id":922,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/896\/revisions\/922"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/897"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=896"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=896"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=896"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}