Get Started
Home
Topics
Search
Library
7 min read · Agents · Code Generation · Added Oct 7 · Paper published Oct 1, 2026

World Editing: Intervening on Executable Worlds at Increasing Depth

Source: research paper via Hugging Face Daily Papers
IGMBench asks whether coding agents can modify existing Minecraft and Terraria worlds via real mod toolchains, not generate or play them. The twist: builds succeed but behavior fails, with top agents dropping from 96% to 48% success as edits couple more systems — compilation is a weak signal for live-system edits.
TL;DR
IGMBench tests whether coding agents can modify existing Minecraft and Terraria worlds through real mod toolchains, finding reliability drops as edits couple more entities, dynamics, and systems, with most failures happening after the mod successfully builds and loads.
Why It Matters
Most AI-and-games research either generates whole worlds from scratch (think video world models producing frames) or trains agents to play inside existing ones. Nobody has carefully studied the middle case: handed an already-working game, can an AI change it in a specific way without breaking everything else? That is what human modders do daily, and it is what you would need if you wanted an AI to produce controlled variants of an environment for training other agents.
The authors from G-G-G, Comfy Org, University of Waterloo, and UIUC frame this as world editing: given an executable world, apply a requested change while preserving unrelated behavior. The relevant baseline isn’t a prior method, it’s just “ask a frontier coding agent to go edit the mod repo.” No specialized world-editing system exists yet to compare against.
How It Works
The contribution is a benchmark and an evaluation harness, not a new model. Two pieces: IGMWorld is the execution environment (prepared mod scaffolds for Fabric on Minecraft and tModLoader on Terraria, a build toolchain, a game server, and hidden validators). IGMBench is 110 human-authored tasks with ~1.1K executable checks written by 11 annotators with modding experience.
The key conceptual move is intervention depth, a four-level axis for how tightly an edit couples pieces of the world:
•
L1 property: tweak an existing value (damage number, drop rate).
•
L2 entity: add new content that conforms to existing rules (a new item).
•
L3 dynamics: introduce new interaction rules or state transitions.
•
L4 system: coordinate multiple components that jointly form a world-level system.
Depth is about semantic coupling, not lines of code. A small L4 edit can still require several interacting pieces to stay consistent.
Each task runs as an edit-build-run-observe loop with a 1-hour wall clock. The agent sees the natural-language request and requirements but never sees the grading logic. After the budget, the submission is scored in three stages:
def evaluate(submission, task): if not builds_and_loads(submission): return FAIL # Stage 1: executability gate for check in task.state_and_behavior_checks: run_check(check) # Stage 2: query game, trigger actions, # compare pre/post, check regressions if task.has_visual_assets: score_visual(submission, factor="semantic") score_visual(submission, factor="art style") # Stage 3 return aggregate(criterion_pass_rate, task_success)
Behavioral checks actually poke the running game: spawn the player, trigger the action, read back entity state or event traces. For stochastic mechanics, they use Monte Carlo sampling, direct parameter inspection, or boundary tests. Visual assets are scored by TPIPS against category-matched vanilla references, with per-category thresholds calibrated from a leave-one-out procedure on the native art.
What They Found
Seven agent configurations were tested, each on each task once: GPT-5.6 Sol and Luna with Codex CLI, Claude Opus 4.8 with Claude Code, Gemini 3.5 Flash with Gemini CLI, plus DeepSeek-V4-Pro, Kimi-K3, and GLM-5.3 under a shared Hermes harness. Two headline metrics: Co-occurrence Prior Regularization counts individual checks passed, World-Editing Success Rate (WSR) counts tasks where all checks pass.
•
Frontier agents are already decent at this. The best configuration (GPT-5.6 Sol) hits 78.2% WSR and 94.8% CPR overall. Weakest sits at 31.8% WSR. The CPR-WSR gap means agents often get most of a task right but miss something.
•
Depth hurts, and not just because deeper tasks have more checks. GPT-5.6 Sol goes from 96.3% WSR at L1 to 48.1% at L4; Claude Opus 4.8 from 92.6% to 55.6%. When the authors stratify tasks by criterion count to control for sheer requirement volume, L1 is still noticeably more reliable than L3-L4. So depth captures something beyond “more things to check.”
•
Most failures happen after the mod loads. Build and load failures are a minority at every level. Behavioral failures dominate, meaning the agent produced something that runs but doesn’t do what was asked. At L1 these are mostly wrong values or formulas. From L2 upward, interaction/progression failures (new content misbehaving against existing systems) grow to 43% at L4, while registration-type failures stay flat at 13-24%. This is consistent with the depth story: deeper edits need coordination, and that is where agents stumble.
•
Visual integration is a separate axis. Native assets jointly pass style and semantic checks at ~0.80. The best agent configurations reach 0.47 in Minecraft and 0.29 in Terraria on the joint check, with everyone below 0.50. The ranking by visual quality does not match the ranking by functional correctness, suggesting functional and perceptual editing are different skills.
•
Agents work differently. GPT and Claude configs inspect local game sources in 91-99% of runs. Gemini leans much harder on web search. All of them rebuild iteratively.
One ablation worth flagging: removing the host-provided generic modding skills (Fabric/tModLoader cheat sheets, a sprite-processing pipeline) dropped GPT-5.6 Luna from 64.5% to 57.3% WSR overall, with the sharpest hit at L4 (63.0% to 33.3%). CPR barely moved. The skills mostly help agents finish all requirements of harder tasks, not get individual checks right.
What’s Useful
•
If you are evaluating a coding agent and your current benchmarks are repo-level patches (SWE-bench-style), this is a different axis: the correctness signal comes from a running game, not a passing unit test. Worth running if you care about agents that modify live systems with runtime behavior you can’t fully specify in code.
•
The depth taxonomy (property / entity / dynamics / system) is portable. If you are designing your own agent eval over a complex executable system (a database, a game engine, a simulator), ordering tasks by how many coupled components an edit touches gives you a cleaner reliability curve than just counting requirements. The paper shows this isolation matters.
•
The build-and-load-then-fail pattern is the actionable finding for anyone shipping coding agents: compilation success is a weak signal. If your agent harness treats “it builds” as near-done, you are probably overestimating success on anything beyond localized edits. Worth adding behavioral post-conditions to your own eval loop.
•
Visual asset generation is an independent failure mode. Even top functional agents produce art that is visibly off from the host style. If your use case involves generated content that must blend into existing assets, do not assume the strongest coding agent is also the best visual choice.
•
The benchmark itself requires owning Minecraft and Terraria and standing up Fabric/tModLoader. The authors release task specs, scaffolds, and validators but not game binaries.
Caveats
•
Each model-task pair runs once. With a 1-hour budget and native game execution, repeats were too expensive. So the numbers are configuration-level snapshots, not variance estimates. Treat small gaps between agents cautiously.
•
Regression checks (does the edit break unrelated things?) exist for only 30 of 110 tasks and are concentrated at L1. Preservation at deeper levels is not well tested, which is a real limitation given depth is the paper’s central axis.
•
Only two games, both with mature modding ecosystems. The authors show qualitative examples in PEAK, Palworld, and Starbound to argue the taxonomy generalizes, but those are not scored.
•
Audio, animation, and narrative are out of scope. “World editing” here means state, behavior, and visual assets.
•
Comparisons are between full agent configurations (model + harness + tools), not model backbones in isolation. A weaker-looking model under Hermes may do better under a provider-tuned harness; the paper cannot separate these.
Topics
Agents
Code Generation
Evaluation
Agents
Code Generation
Evaluation
Up next in Agents
Optimizing the Optimizer: Language Models Discover Faster Molecular Relaxation
Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents231 episodes
Evaluation167 episodes
Code Generation51 episodes