Trace2Env turns recorded agent interaction logs into a structured worldbook that a world-model agent consults, with explicit per-episode state, so simulated terminals and web apps stay consistent across turns and replayed actions remain valid in the real environment (CR 0.914 vs 0.032 on ALFWorld).
Imagine you want to train or stress-test an LLM agent against a real terminal, enterprise portal, or web app. Standing the real system back up for every experiment is often impossible: you lack the source, the data snapshot, or the infrastructure. What you often do have is logs of past agent sessions: the commands issued, the screens returned, the errors that followed.
The standard workaround is a Language World Model: feed a description of the environment plus the interaction history into an LLM and ask it to predict the next observation. This breaks on long horizons. If the agent deletes report.txt at turn 3, nothing in the immediate output flags that fact, and by turn 14 a cat report.txt call might cheerfully return invented contents. The simulated world drifts away from any real one, which means success inside the sim does not predict success outside it.
Trace2Env splits the job in two. Offline, it reads the historical traces once and compiles a reusable worldbook describing the environment. Online, a separate world-model agent acts as the environment: for every action the task agent takes, it decides what state changed and what text to return back.
The worldbook has four parts: action and state schemas (what the simulator can represent), grounded evidence (recorded transitions kept verbatim, so the simulator can see the real output format), induced abstractions (rules, invariants, response contracts distilled across traces), and provenance linking abstractions to the turns that support them. Rules are only promoted to “supported” if they survive review against contrastive examples; otherwise they stay tentative.
At runtime the agent sees a workspace with three distinct scopes: the immutable worldbook (how the environment behaves in general), the mutable episode state (what is true in this run), and episodic memory (what already happened this run). A key move: retrieved worldbook entries pass through an applicability gate labeling them supporting, uncertain, or format-only. Format-only items contribute output shape but not facts, so the simulator does not smuggle in file names from a different session.
The agent proposes a transition. A separate harness validates it against schemas and constraints before committing.
def step(action, workspace):
brief = build_brief(action, workspace) # state + relevant worldbook
while agent.needs_more_info() and budget_left():
agent.inspect(workspace) # rules, evidence, memory, state
proposal = agent.submit(effects, observation, rule_ids)
if harness.validate(proposal, workspace):
state, obs = harness.commit(proposal)
return state, obs
return fallback(proposal)
Across nine environments (terminal, SWE, Android, web, three EnvScaler business apps, and two text games), Trace2Env was evaluated two ways: single-step next-observation fidelity scored by a GPT judge on five dimensions, and multi-turn consistency.
•
On single-step prediction with the GPT-5.6-sol backbone, Trace2Env averages 75.94 vs 69.51 for direct prompting, and best in 11 of 14 environment/backbone pairs. Factuality and consistency improve in every pair.
•
Ablations separate the pieces: handing the worldbook to the model as plain prompt context (Worldbook Prompting) already beats raw retrieval (Trace RAG), and the full agentic runtime adds further gains. Neither component alone explains the result.
•
The multi-turn finding is the sharper one. On ALFWorld, direct prompting shows 97% success inside the simulator but only 3% when those action sequences are replayed in the real environment. Trace2Env shows 88% simulated success and 85% on replay, a consistency ratio of 0.914 vs 0.032. On SciWorld, consistency ratio is 0.706 vs 0.529.
•
The authors’ explanation: direct prompting hallucinates “successful” transitions and then stays internally coherent inside the resulting fiction. Trace2Env rejects unsupported actions as no-ops and preserves that outcome, so the task agent sees honest failure and can recover.
•
Gains persist on evaluation tasks that do not overlap the construction traces, suggesting the worldbook captures reusable environment behavior rather than memorizing task-specific transitions.
•
If you are building offline eval harnesses for an agent whose target system you cannot run (an enterprise portal, a locked SaaS tool), this paper argues that trace-reconstructed simulation is a viable substitute, and more importantly that in-simulator success is a misleading metric. Measure whether sim-generated action sequences still succeed when replayed on the real system; the paper calls this W2R and its ratio to real-env success is the quantity worth tracking.
•
If you are already using a prompt-based LWM, the ablations suggest the cheapest upgrade path is structuring your traces into schemas plus retained evidence turns, even without the full agentic runtime. On terminal, that alone lifts the judge score from 55.76 to 59.07 at one model call per prediction. Adding the agent loop costs roughly 5x calls for another few points.
•
Worth testing: whether the same worldbook built by one model can be operated by a cheaper one. The paper reports that worldbooks built with GPT also worked when operated by DeepSeek-V4.1-Flash, but does not test the reverse direction.
Fidelity is bounded by what the traces show. Behaviors never exercised in logs cannot be reconstructed, and the worldbook does not try to invent them. The paper is honest that the construction-trace scaling curve on terminal is still rising at 36 traces, so returns likely improve with more data at proportional offline cost. Judge scores depend on a GPT-5.2 LLM-as-judge protocol inherited from AgentWorldBench, with its own calibration quirks. The strongest multi-turn result (ALFWorld, CR 0.914) comes from a text game with a small, well-defined action set; whether the same gap holds on messier environments like real enterprise web apps is not established, though the single-step gains do extend there. Finally, Trace2Env is explicitly learning-free: it moves cost from training to inference-time agent loops, which may not fit latency-sensitive uses.