HexaAnything proposes Physical Coding: robots keep their world state and their behavior as inspectable executable programs, so a separate verifier can check each step against fresh evidence instead of trusting a Vision-Language-Action model’s stop signal, and failed steps become localized code edits rather than whole-episode do-overs.
Today’s dominant recipe for foundation-model-driven robots is to train a Vision-Language-Action model or a World-Action Model that maps camera frames plus a language instruction directly to a chunk of joint actions, then a stop signal. The paper points out an awkward property of this setup: on the popular LIBERO benchmark, a policy trained with the instruction masked out still succeeds 92.3% of the time, versus 96.2% with the instruction. The model is mostly learning “given this scene, execute this trajectory,” not “follow this instruction.” It also collapses under small viewpoint or object-layout changes.
The deeper diagnosis: an action chunk plus a stop token has no slot for which subgoals are done, which preconditions hold, or what to do after a failure. If a controller closes the oven door with only one ingredient inside, nothing in the interface can detect or repair that. Prior code-driven robot work like Code as Policies moved the policy into code, but trusted perception outputs and stored improvements outside the model in skill libraries. Embodied harnesses like Zetta and SHAPER evolve critics and skills around a frozen reasoning model.
The contribution is an interface, not a new policy architecture. Two coupled executable artifacts run per task:
•
Code as World: a program whose entries record objects, relations, measurements, constraints, and boolean progress predicates (e.g. “both ingredients inside oven”). Each entry carries provenance (which tool call, which timestamp, which experimental condition). It’s built from images in the style of VCode, not from privileged simulator state.
•
Code as Policy: a program with typed nodes for observation, action, tool call, verification, branching, and recovery. Conditions are predicates of the world program.
The Harness is the runtime around these two artifacts. After every tool call, it re-observes, updates the world program, and runs an independent verifier that returns typed verdicts: Pass, Fail, InsufficientEvidence, Blocked, or SafetyStop. Crucially, a model’s own “I’m done” claim is never accepted as success, because verified traces later become training data and a self-scoring verifier would confirm its own mistakes. A Vision-Language-Action model is demoted to one available action tool that the policy program can call, interrupt, or replace.
while not goal_predicates_hold(W, evidence):
node = policy_model(task, W, P, history)
outcome = dispatch(node, tools) # may call a VLA
obs = observe()
W = update_world(W, obs, outcome)
verdict = verifier(W, outcome, task)
if verdict in {Fail, InsufficientEvidence}:
P = recover_or_replan(P, W, verdict)
elif verdict in {Blocked, SafetyStop}:
break
Every step’s world entries, policy node, tool outcome, and verdict get logged. A failure can then be traced to a specific predicate, tool, or recovery branch and edited there. Verified successful traces are admitted (after static checks and regression tests) as training data for the next model, called HexaModel.
Three pieces of evidence, all with the same Vision-Language-Action model held fixed where relevant.
•
Harness helps a fixed Vision-Language-Action model. On RoboCasa365, plugging XR-1 into HexaAnything (driven by GPT-5.6-Sol) lifts Composite-Unseen success from 34.3% to 38.3%, Composite-Seen from 54.8% to 61.5%, and overall from 56.6% to 61.1%. A generic Codex coding agent with the same underlying model only reaches 59.5% overall and doesn’t beat native XR-1 on Composite-Unseen, suggesting the typed workflow/verifier contracts matter, not just “wrap it in a coding agent.” On three specific Composite-Unseen tasks with 100 seeds each, gains are +31, +15, +17 points.
•
Traces train a better planner. HexaModel v0.1 is Qwen3.8-27B fine-tuned on 9.3K VQA examples plus 1K harness traces (from GPT-5.6-Sol, the base Qwen model, and an earlier checkpoint). Placed back in the same harness, it beats its base on every split and edges out GPT-5.6-Sol overall (61.7% vs 61.1%), with the largest gain on Composite-Unseen (the split most about decomposition and recovery).
•
Tools can be revised against a fixed evaluator. On three RoboDojo tasks, a programming agent reads failed traces and rewrites tool code. Over two rounds, Fold cloth, Pour vase, and Press by number go from 0/40/0% to 80/100/100%. Improvement isn’t monotonic (Pour vase drops to 0% at round 1 before recovering). The weights never change. On Fold cloth specifically, the authors flag that the human operator had seen privileged garment keypoints; a re-evaluation without that information gives 3/5 on development seeds and 4/5 on held-out seeds.
On PhyBench, a simulated lab, HexaAnything with Opus 5.5 or GPT-6-Astra autonomously plans, runs, and analyzes physics experiments (Hooke’s law, pendulum g, coupled oscillators), reporting mean relative error below 5% on all three. On a real dual-arm AgileX robot, five of seven tabletop tasks succeed in all three trials, with speed and token-count improvements over published references (which used different hardware, so the comparison is loose).
•
If you’re building an LLM-orchestrated robot stack on top of a frozen Vision-Language-Action model: the concrete lesson is to never treat the Vision-Language-Action model’s termination signal as task success. Add an independent state check between action-model calls, gated on a predicate written against fresh observations. The paper’s evidence is that this alone (same weights, same action budget) can add ~15–30 points on long-horizon composite tasks. Doubling the Vision-Language-Action model’s action budget in the native condition did not help.
•
If you’re planning to fine-tune a planner from agent traces: the paper’s model-evolution story only shows a small overall gain (+1.2 pp overall, +2.2 pp on the hardest split) after training on ~1K traces plus VQA data. Worth trying if you already run a harness, but don’t expect the harness gain to be recovered by the model alone. Also note: the verifier that labels those traces must be independent of the model being trained, or you’ll bake in its errors.
•
If you want the tool-evolution loop: it needs a fixed, independent evaluator to score revisions against, plus static checks and regression tests. On real hardware the authors substitute operator judgment and execution traces (measured joint speeds, gripper state, where things landed) for a simulator evaluator. Improvement is not monotonic, so keep versioned rollbacks.
•
For scientific-experiment automation: PhyBench is worth knowing about as a benchmark where success is a numerical claim backed by evidence, not object arrangement. Sequence-of-actions policies structurally can’t do this; you need somewhere to record measurements with their conditions and fit models to them.
•
The headline harness comparison uses one action model (XR-1) on one benchmark (RoboCasa365), and HexaAnything’s overall margin over a generic Codex agent driven by the same LLM is only 1.6 points. Codex is actually ahead on Atomic-Seen. So “typed workflow beats generic coding agent” is directional, not decisive.
•
The RoboDojo tool-evolution numbers rest on five seeds per task. On Fold cloth, a separate programming agent revised the tool while a human operator chose pixels for each call, initially with access to privileged keypoints. The authors’ honest re-evaluation without that access (3/5 dev, 4/5 held-out) is the number to remember, not the 80% in the main table.
•
Real-robot results are three trials per task. Speedup and token-count comparisons are to published numbers on different hardware, not head-to-head.
•
“Self-evolution” here means Harness edits, tool edits, and one round of retraining a planner from verified traces. It does not mean autonomous architecture search, weight self-modification, or unconstrained deployment. The authors explicitly flag this.
•
Because verified traces train the next model, the verifier and the planner must not co-adapt. The paper names this as an open problem, not a solved one.