Get Started
Home
Topics
Search
Library
6 min read · Inference Optimization · Robotics · Added Sep 29 · Paper published Sep 25, 2026

InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

Source: research paper via Hugging Face Daily Papers
0:00 / 7:55
InternW0-Δ tackles the cost of world-model policies: predicting future video helps robots plan but is too slow at control time. The fix is Causal Imprint—future frames supervise dedicated tokens only through gradients, masked from forward attention, so the future branch is deleted at inference, hitting 92.8% on LIBERO-Plus.
TL;DR
InternW0-Δ is a robot manipulation model that uses future video frames as a training-only teacher signal, so at deployment the action policy inherits predictive scene knowledge without having to actually generate future video. On the LIBERO-Plus robustness benchmark it reaches 92.8% success under distribution shift.
Why It Matters
Suppose you’re training a robot to pour water or toast bread from camera input. One popular recipe, called a World-Action Model, is to train a single model to both predict what the scene will look like next AND predict the next robot action, on the theory that being able to imagine “the cup tips and water flows” should help you plan the pour. The problem: predicting pixels is expensive, and at control time you don’t want to burn compute rolling out imaginary video before every action. You want actions, fast.
The usual workaround in the Vision-Language-Action model family (models like π₀ or OpenVLA) is to skip world modeling entirely and just map vision+language directly to actions. You lose the predictive prior. Recent WAMs like Fast-WAM separate future-video supervision from the action inference path, and InternW0-Δ builds on that direction.
How It Works
The architecture couples two Transformer stacks that share attention but keep separate weights, a pattern the paper calls a Mixture-of-Transformers. One stack (“video expert”, initialized from a pretrained Wan2.2 video model) processes camera frames. The other (“action expert”, trained from scratch) produces the actual robot commands. A frozen vision-language model, RynnBrain1.1-2B, feeds scene understanding into the action expert so it knows what objects are where.
The key trick is called Causal Imprint. The authors add learnable tokens inside the video expert whose job is to encode “how is this scene about to change?” During training these tokens are supervised by actual future video latents (differences between consecutive future frames, plus a cosine alignment to deeper future features). But the attention mask forbids the imprint tokens and the action expert from ever reading the future directly. Future frames only shape gradients, never activations. At inference the future branch is deleted entirely, so action generation is a single video-encoder pass plus a small diffusion loop on actions.
A second training-only signal comes from Track4World, a frozen model that produces geometry and motion descriptors. The video expert is trained to reproduce these via an MSE loss (Feature distillation), again with the teacher discarded at inference.
# training step (schematic) z_past, z_future = vae(frames_past), vae(frames_future) h_video, h_imprint = video_expert(z_past, imprint_tokens) # attention mask: h_imprint cannot see z_future in forward pass action_pred = action_expert(h_video, h_imprint, vlm_context, state) loss = (fm_loss(action_pred, actions_gt) + fm_loss(video_pred, z_future) # future supervision + mse(h_imprint, delta(z_future)) # Causal Imprint target + cos_align(h_imprint, h_video_future) # alignment to future feats + mse(student_desc, track4world_desc)) # 4D distillation # inference: no future branch, no distillation branch action = action_expert(video_expert(z_now, imprint_tokens), vlm_ctx, state)
Training data spans 20K+ hours from 15 robot datasets plus egocentric human video (EgoDex, EgoVerse), UMI (Universal Manipulation Interface) hand-held grippers, and human videos converted to robot demonstrations via an Ego2Robot pipeline. Everything is remapped into a canonical 80-dimensional action vector with fixed semantic slots (arm joints, end-effector pose, gripper, hand, torso, base), so different embodiments share the same output space.
What They Found
All simulation numbers are with the same pretrained checkpoint, post-trained per benchmark:
•
LIBERO-Plus (robustness to camera/lighting/language/layout shifts, trained on clean LIBERO only): 92.8% overall, versus 91.4% for the strongest VLA baseline (Qwen-RobotManip-Context) and 84.8% for the strongest prior WAM (Being-H0.7). Biggest jump is under robot perturbations: 91.1% versus 87.4% prior best.
•
RoboTwin 2.0 Clean2Random (trained clean, tested on randomized scenes): 71.9% versus 69.4% best prior; also leads Clean2Clean at 90.0%.
•
EBench (mobile bimanual): overall score 66.0, with the largest advantage on long-horizon tasks (score 76.5).
•
RoboDojo: average success 23.9%, roughly double OpenWAM-α (11.9%). Still low in absolute terms because RoboDojo is deliberately hard.
The cumulative ablation on LIBERO-Plus is the most informative piece. Starting from a 49.6% baseline: adding sparse memory context (+3.9pp), swapping in the RynnBrain VLM (+15.6pp over baseline+SMC), adding Causal Imprint’s difference loss (+1.7pp), then its alignment loss (+5.7pp), then 4D distillation (+1.9pp), reaching 78.4%. The VLM choice and the Causal Imprint alignment objective are doing the heaviest lifting. Track4World beats other geometry teachers (CoWTracker, Pi3X) tested.
Data ablations: robot demonstration data helps most (LIBERO-Plus 78.6% → 83.7%; RoboTwin Clean2Random 4.4% → 32.3%). UMI data helps meaningfully (RoboTwin Clean2Random → 13.9%). Raw egocentric human video barely moves success rate on these gripper benchmarks; the Ego2Robot conversion pipeline recovers some of that value.
Real-robot deployment on four platforms (two gripper, two dexterous-hand) shows pretraining matters concretely: on “toast bread” success went from 4/20 to 19/20 with pretraining; on a Luminol chemistry task, 0/20 to 19/20. Inference latency on one RTX 5090 was optimized from 780ms to 152.8ms round-trip (5.11× speedup), which is what enables 30 Hz control with asynchronous action chunks under Real-Time Chunking.
What’s Useful
If you’re building a WAM-style policy and worried about the cost of generating video at control time, the Causal Imprint pattern is a concrete alternative worth testing: use future frames only as a loss target on dedicated tokens, mask them out of forward attention to the action head, and delete the future branch at inference. The ablation isolates this contribution and shows it accounts for ~7pp on LIBERO-Plus, which is a bigger swing than the 4D distillation piece.
If you’re pretraining across heterogeneous robot datasets, the canonical 80-dim action layout with fixed semantic slots (plus per-embodiment validity masks) is a practical unification recipe. The paper releases the filtering pipeline and versioned data indices, which matters more than the numeric hyperparameters.
One caveat before copying the data mix: on these gripper-heavy benchmarks, raw egocentric human video contributed little. The authors explicitly note this may not transfer to dexterous-hand tasks where fingertip trajectories are more informative, and their evaluation only measures success rate, not representation quality. Don’t over-generalize “human video doesn’t help” from this evidence.
For async execution on real hardware, they report that training-time prefix conditioning (from the Real-Time Chunking recipe) produced much smoother chunk-boundary transitions than either action blending or the VJP guidance approach (boundary-to-within-chunk change ratio of 1.12 versus 7-9). If you’re stitching action chunks, this is worth the training-time investment.
Code, checkpoints, data-processing tools, and deployment utilities are promised but the paper doesn’t link a specific repo in the excerpt; the project page is the entry point.
Caveats
The 5.11× inference speedup is on a specific dexterous-hand deployment on one RTX 5090, not a general claim. The RoboDojo numbers are still low in absolute terms (23.9% average success); memory, precision, and open-vocabulary categories remain hard for everyone including this model. Baseline numbers on all four benchmarks are taken from other papers under matching protocols, not re-run by the authors, so exact head-to-head reproduction may differ. Finally, the GPT-guided policy exploration in Section 7.5 is a small (5-task, low-episode) probe, not a systematic study; treat the 10.8% → 47.2% jump as suggestive rather than a reliable estimate of what agent-assisted control buys you.
Topics
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Inference Optimization107 episodes
Robotics54 episodes
Video Generation62 episodes