EditHero is a 3D editing benchmark where each object goes through up to 30 part-level edits in a row, with an exact rebuilt reference after every turn, exposing that today’s native-3D editors drift badly while LLM/VLM agents preserve untouched parts much better but take minutes per edit.
3D artists almost never make one edit and stop. A character gets a hat, then loses the hat, then gets different hair, then a new outfit. Each revision has to keep everything the earlier revisions established. The industry calls the aspirational version of this vibe modeling: shaping an asset through back-and-forth natural-language conversation.
The problem is that current text- and image-guided 3D editors are evaluated on a single edit from a clean source. Nobody has measured what happens when you ask them to edit their own output five or ten times in a row. Prior paired datasets only give you isolated before/after pairs, and their targets are often themselves generated by another editing model, so you can’t tell whether an edit preserved the untouched parts or quietly rewrote them.
EditHero fills this gap. It gives you 457 edit chains (2755 total edits across 252 objects, median 6 turns, up to 30) where the exact correct 3D state after every turn is known by construction, so you can score each turn fairly.
The key move: instead of generating target states with another model, the authors assemble them from a library of parts. Each starting object (a “host”) is pre-segmented into named slots like “head”, “porch canopy”, “saddle”. A data engine then applies one of four atomic operations per turn: add (attach a library part to a free contact face), remove (delete the part in a named slot), replace (swap a part for another at the same mount), or retexture (restyle a slot’s appearance while keeping its geometry).
Each accepted operation is logged to a JSON file along with the chosen library part, pose, and texture. The log plus the part library deterministically rebuilds any intermediate state, so there is an exact ground-truth mesh after every turn. Retextures use Qwen-Image-Edit to restyle a slot render, then TRELLIS.2’s texturing module paints the unchanged geometry. Every chain was then human-reviewed and repaired turn by turn.
For evaluation, methods run under self-rollout: at turn k, the input is the method’s own output from turn k-1, not the ground-truth previous state. This is the realistic setting, and it’s where errors compound.
state = source_mesh
for turn in chain:
instruction, target_render = turn.prompt, turn.view
state = method.edit(state, instruction, target_render)
score_edit_region(state, turn.gt) # instruction following
score_unchanged_region(state, prev_output) # content consistency
prev_output = state
Two custom metrics avoid the usual pitfalls. Instruction Following (IF) is measured only in the region defined by the named parts (before and after the edit), and it’s computed against the method’s own previous output, not the ground-truth previous state. Otherwise a no-op (returning the input unchanged) would get credit for removals it never made. Content Consistency (CC) measures how much of the untouched region survives, either compared to the previous output (CC-prev) or the original host (CC-0). Two reference points anchor the scale: a no-op baseline (IF=0, CC=1 by construction) and a regeneration baseline that runs TRELLIS fresh on each target render (high IF, low CC).
Four native-3D editors were tested: PartFlow, Nano3D, 3DEditFormer, and VoxHammer. All regenerate the whole object through a learned 3D latent, so every turn risks changing every corner.
Whole-object scores reward doing nothing. On whole-object F-score, LPIPS, and PSNR, the no-op baseline beats every non-agentic method. A single edit touches only a small region, so returning the input scores well. This is why region-split metrics matter.
Instructions are followed poorly from turn 1. IF ranges from 0.12 to 0.32 over all turns, and is only 0.24 to 0.36 at turn 1, before any accumulation can happen. Nano3D and VoxHammer follow instructions best, but VoxHammer keeps much less of the rest of the object (unchanged-region IoU 0.71 vs 0.89).
Errors compound under self-rollout. From turn 1 to turn 7, CC-0 (how much of the original host survives in regions never asked to change) falls from 0.60-0.78 down to 0.33-0.67. A reset diagnostic confirms much of this is inherited: when turn k starts from the ground-truth state instead of the method’s own output, IoU to target at turn 5 is 0.68 for Nano3D vs 0.52 under self-rollout (and 0.46 vs 0.36 for 3DEditFormer).
LLM/VLM agents behave opposite. Six models were tested on a 55-chain comparison subset: GLM 5.3 Flash, DeepSeek V4.1 Flash, Sol 6, Astra, Fable 5.1, and Opus 5.5. Agents inspect the current mesh, write code to modify only the named part, and render-and-compare before committing. By construction they leave the rest untouched. All six agents keep the unchanged region better than any non-agentic method (CC-prev 0.95-0.98 vs at most 0.90). Five of six also follow instructions better (IF up to 0.62 for Opus 5.5 vs at most 0.36). Removal is easy for agents (IF 0.88-0.99, since code just deletes). Addition and replacement are harder (IF 0.10-0.49 and 0.17-0.52) because code-built geometry often misses the target’s shape.
Agents are slow. An agent takes 1.5-6 minutes and 6-16 model calls per edit (medians). Non-agentic methods take 20-50 seconds on a local GPU.
For practitioners deciding how to build an iterative 3D editing tool, the results point in a specific direction:
•
If you’re shipping an interactive editor where responsiveness matters, the current native-3D editors are fast but will quietly degrade your asset over a handful of turns. Measure your own system on a self-rollout protocol before trusting single-turn benchmark numbers. The whole-object F-score you’ve been reporting may be meaningless if a no-op beats you.
•
If quality of preservation matters more than latency (batch tooling, revision workflows, offline asset refinement), the agentic path is already competitive and preserves untouched regions far better. The bottleneck isn’t the approach, it’s that code-built geometry is crude. Worth testing whether you can hand the agent a learned part-generator as a tool rather than forcing primitives-only.
•
The authors themselves flag the open question: can you combine agentic control (local, preserves state) with learned 3D generators (produces good-looking geometry fast)? The benchmark exists to measure such hybrids.
•
When evaluating any multi-turn editor, split metrics by edit region and unchanged region, compare against the method’s own previous output rather than ground truth, and always report the no-op baseline. Otherwise “do nothing” wins.
Code and data are released: GitHub, project page.
•
Targets come from a part library plus deterministic placement, so the geometry an “ideal” method would produce is library-shaped. Methods that invent new, equally valid geometry may score low on IoU even when the result is good.
•
The reset diagnostic covers only Nano3D and 3DEditFormer on 200 chains, so the “errors are inherited” conclusion is strongest for those two.
•
Nano3D has no native material mode; its retexture turns are routed through its replace mode, which may understate it on material edits.
•
Agent comparisons use a 55-chain subset (416 turns), not the full 457 chains, because of cost. The specific agent models tested (Opus 5.5, Astra, Fable 5.1, etc.) are the authors’ choices; the paper does not fully specify runtime configurations beyond what’s in its appendix.
•
The paper evaluates editing fidelity, not artistic quality. An agent that keeps the asset intact but builds ugly new parts still scores well on CC.