AgentGarten builds interactive training worlds by pairing a code-defined physics engine with a shared neural video renderer distilled via Adversarial Forcing, letting pretrained agents learn tool use in a handful of rounds instead of millions of RL episodes.
If you want an LLM-based agent to practice an embodied skill (navigate a room, drive over a bridge, herd sheep), you need an environment that is two things at once. It must be faithful: state, physics, and rules behave consistently so the agent’s actions have reliable consequences. It must also be realistic: the pixels the agent sees look like the real world, so a model pretrained on web images can actually interpret them.
Today you pick one side. Classical simulators like Habitat 2.0 or ManiSkill2 give you clean programmable state, but every new scene needs hand-made 3D assets, textures, and lighting. Video world models like Genie or GameNGen produce rich-looking frames, but the “state” lives implicitly inside generated history, so you cannot inspect it, edit it, or guarantee that pressing a button yesterday still matters tomorrow. The paper’s bet is that you can split the job cleanly: an engine owns state, a learned renderer owns appearance.
The core idea is a two-box architecture. A scene program runs in a normal simulator and maintains ground-truth state: object poses, articulations, task variables, camera. At each step it exports a structured condition, specifically a colorized depth map or surface-normal map, which is just the raw geometry of what the camera sees. A shared neural renderer then paints that geometry into a photorealistic frame, using an appearance reference image, a text caption, and its own cached visual history to fill in colors, materials, and style.
Crucially, the engine never reads a rendered pixel back. The renderer only affects state through the agent’s action choices, so visual hallucinations cannot corrupt physics. You can even re-render the same recorded rollout in a different visual style without changing what happened.
The renderer itself is a Cosmos 3-Nano video model adapted in three stages. First, it is retrained to accept geometry tokens alongside RGB tokens through joint attention (no pixel-aligned control branch like ControlNet adapter). Second, it is converted from bidirectional to block-causal generation: it emits short blocks of a few latent frames at a time so an agent gets fast feedback. Third, it is distilled with Adversarial Forcing, the paper’s main technical contribution, into a few-step real-time renderer.
Adversarial Forcing extends Self-Forcing with two fixes. The first is exact replay. Standard self-forcing distillation runs a rollout, caches past keys/values with gradients detached, then computes losses on later blocks. Gradients never flow back into how history was encoded. A prior method, Self Gradient Forcing (SGF), adds a second differentiable pass but computes it with a different attention kernel, so numerical results drift from the sampled rollout. AgentGarten instead replays the rollout block-by-block using the exact same attention calls and tensor shapes, producing a bitwise-identical recomputation that still carries gradients through history. The second fix adds a GAN loss against real videos (in the style of DMD2 (Distribution Matching Distillation 2) and R3GAN) to prevent the texture degradation that pure distribution-matching distillation shows over long rollouts. Because the discriminator’s R1/R2 regularization normally requires double-backward through fused attention kernels (which FlashAttention does not support), the authors derive an exact penalty computed from one vector-Jacobian product plus one Jacobian-vector product through the frozen backbone, with no backward-through-backward anywhere.
# One Adversarial Forcing training step (conceptual)
with torch.no_grad():
rollout = []
for block in range(num_blocks):
u, z = student.sample_block(context, cache) # record input u, clean output z
cache.append_detached(z)
rollout.append((u, z))
# Replay: same per-block attention calls, but differentiable
for j, (u, z) in enumerate(rollout):
z_pred = student.predict(u.detach(), history=[r[1].detach() for r in rollout[:j]])
loss_dmd = dmd_loss(z_pred, teacher, fake_score)
loss_gan = adversarial_loss(z_pred, real_clip, discriminator) # exact R1/R2
(loss_dmd + lam_g * loss_gan).backward()
On top of this, agents live in a round-based practice loop. Each round, the agent gets a task file (goal, actions, limits, no solution), plays with only rendered first-person frames, then writes playbooks: markdown skill files summarizing what worked. The next round’s agent starts in a fresh conversation but inherits the accumulated playbooks.
On visual quality, the paper shows 30-second rollouts where plain Self Forcing develops repetitive surface patterns and loses detail, while Adversarial Forcing holds texture throughout. This is a qualitative figure, not a FID number.
On exact replay, their block-by-block SDPA execution matches the sampled rollout bitwise, while a full-sequence FlexAttention replay (the SGF-style approach) diverges by 3.99% relative L2 error. Forward time drops 5.4% with negligible memory change. The claim here is numerical fidelity of the gradient path, not a downstream accuracy gain.
On speed, the renderer produces 480x832 video at over 35 FPS on one H100, counting condition encoding, four denoising steps, decoding, and host transfer, using custom Triton kernels, CUDA graphs, and a tiny distilled VAE decoder.
The headline agent result revisits Hide-and-seek (Baker et al. 2019), OpenAI’s 2019 multi-agent physics benchmark where RL from scratch needed roughly 25 million episodes before hiders built shelters and ~100 million before seekers used ramps. With a pretrained foundation-model agent (the paper cites a GPT-6 variant) that sees only rendered first-person frames and keeps playbooks, hiders built panel shelters by round 4 and seekers crossed walls with ramps by round 10. The authors are careful: they frame this as a different learning paradigm, not a sample-efficiency ratio, because the pretrained agent already knows what a ramp is and only needs to ground that knowledge in closed-loop action.
Four additional worlds (companion dog, one-lane bridge, herding, quarry loader) show the same round-over-round improvement pattern. For example, bridge-swap time fell from 71s to 41s between round 1 and round 4; the quarry loader went from delivering nothing to clearing rocks, delivering one, and parking by round 4. These are single-episode-per-round numbers in each world, not statistical comparisons against a baseline.
If you are building an agent evaluation harness and your pain point is asset creation for visual diversity, the architectural pattern here is worth copying even without reimplementing the renderer: keep state in a programmable engine, let a learned model handle appearance, and expose a geometry-only interface (depth or normals) between them. This lets a coding agent spin up new scenes from a text prompt or image without authoring textures.
If you are doing autoregressive video distillation, the exact-replay trick is a drop-in improvement over SGF-style second passes whenever you care about gradient consistency with the sampled trajectory. The exact R1/R2 construction is also directly reusable for anyone training a GAN discriminator on top of a frozen fused-attention backbone, where double-backward is blocked.
If you are evaluating agents that learn from written experience (in the lineage of Reflexion or Voyager), the playbook-across-fresh-conversations setup is a cleaner test of whether lessons generalize than keeping context in-session. Worth testing on your own tasks: does an agent that must write its lessons for a stranger write better lessons?
Be careful about one scope issue. The paper demonstrates emergent tool use with a specific pretrained agent across a handful of hand-authored worlds. It does not show that playbook-based practice beats fine-tuning, nor that the renderer is good enough for sim-to-real transfer. Treat the hide-and-seek numbers as an existence proof for grounding priors, not as a benchmark score.
Geometry-only conditioning cannot express state that leaves depth and normals unchanged: a spinning symmetric wheel, a color change, a material swap. The authors flag this as an open problem. Visual history carries some of it, but only within the memory window.
The agent comparisons to the 2019 hide-and-seek RL results are not a sample-efficiency claim and the paper says so. Pretrained agents start with web-scale priors; RL started from random weights on privileged state. The two numbers measure different things.
The four additional worlds report one episode per round, so round-over-round improvement is suggestive rather than statistically established. The paper does not run ablations isolating which component (renderer quality, playbook inheritance, pretrained agent) drives the gains.
Finally, the renderer runs at 35 FPS on a single H100. Interactive deployment at this fidelity assumes datacenter-class hardware per agent session.