Long-WAM lets a robot condition its next move on up to ~19 seconds of past video and an imagined short-term future, by first pretraining the video backbone Answer Refusal (AR) on robot and egocentric footage and only then adapting it to output actions. Longer history helps far more when pretraining is autoregressive than bidirectional.
If you build a robot policy that eats camera frames and spits out joint commands, a single current frame tells you where the mug is but not that it is sliding toward the table edge. Policies that only see the current frame miss motion cues and lose track of multi-stage tasks.
The natural fix is to feed in a window of past frames. But past work on World-Action Model shows two problems. First, longer history means more tokens, more attention, more latency, and robots have to act now. Second, and this is the paper’s main empirical point, just letting the model see more history does not mean it uses it. A video backbone pretrained bidirectionally (like most text-to-video models, where every frame can attend to every other) does not reliably improve when you extend its history window after fine-tuning it for control. The authors contrast this with autoregressive pretraining, where the model is trained to predict future frames from past ones, which matches the causal structure a controller actually needs.
The practical baseline here is a class of recent video-conditioned policies (LingBot-VA, DreamZero, Fast-WAM) that adapt bidirectional video generators into causal controllers. Long-WAM instead keeps the pretraining itself causal.
The pipeline has three parts.
Stage 1, causal video pretraining on robot footage. The authors continue training an existing long-video autoregressive model (LongLive-2.0) on roughly 10,000 window-equivalent hours of robot and egocentric video, with no action labels. The model learns to predict the next chunk of video latents given earlier chunks, using teacher-forced Flow matching with a trick they call error recycling: during training some context frames are perturbed with the model’s own past prediction errors, so the model learns to recover from its own mistakes during long rollouts.
Stage 2, action adaptation while preserving causal structure. They attach a second expert (an action head) alongside the video model. At decision time: the video expert partially denoises a few future Variational Autoencoder latents from the observed history, and the action expert then denoises an action chunk while attending to both observed history and those predicted-but-not-fully-clean futures. Crucially, the predicted future frames are never decoded to pixels, they stay as cached keys and values. The authors call this inverse dynamics modeling (IDM), as opposed to co-denoising video and actions jointly.
# at each decision step t
obs_latents = streaming_vae(recent_frames) # encoded as frames arrive
future_latents = video_expert.rollout( # partial denoise, stop at sigma=0.9
noise, context=obs_latents, steps=4)
kv_cache = video_expert.prefill(obs_latents, future_latents)
action_chunk = action_expert.denoise( # 4 steps, attends to kv_cache
noise, kv_cache, robot_state, instruction)
execute(action_chunk[:R]) # overlap with next inference
Stage 3, deployment system. To keep this real-time, they run inference asynchronously with robot execution, encode incoming camera frames through a streaming causal Variational Autoencoder so the prefix is ready before the inference trigger, and quantize the video expert’s linear layers to NVFP4. Action compute stays in BF16 for precision. They add CUDA Graph replay replay and kernel tuning per device.
Pretraining style determines whether history helps. On RoboCasa GR-1 Tabletop, extending history from 0 to 19.2 seconds raises success from 63.3% to 78.7% with their robot-domain autoregressive backbone. A bidirectional backbone (Wan2.2) under the same causal adaptation goes from 61.7% at 2.4s to 64.1% at 9.6s and back to 61.6% at 19.2s. The gap between the two backbones grows from 3.3 points at zero history to 17.1 points at 19.2s. On LIBERO-Long a shorter 2.4s window suffices, moving success from 94.5% to 99.5%.
Diminishing returns at very long windows. At 38.4 seconds, success drops to around 75%. The authors attribute this to padding: in their training trajectories, 80.4% of the sampled history frames at that window are repeats of the first frame rather than real history.
Predicting future video before acting helps on long tasks. The IDM denoising strategy beats history-only (w/o V) and joint video-action co-denoising (CoD) by the largest margin on LIBERO-Long: 99.5% vs 94.5% vs 97.8%. The predicted future latents are never rendered to pixels; they only serve as a cached conditioning signal.
Real-time latency on edge GPUs. On an RTX 5090 each action chunk takes 107.4 ms including the full observation encoder, a 3.3x speedup over the unoptimized BF16 baseline. The system also runs on DGX Spark and Jetson AGX Thor in the 328-379 ms range.
Real robots. On a Unitree G1 grasping cups from a conveyor at 7.5 cm/s, Long-WAM succeeds in 90% of 20 trials, and in 95% on a dynamic cup-stacking task, where two baselines (Fast-WAM and \pi_{0.5}) succeed in zero of 20 trials each. On a bimanual YAM platform, three long-horizon tasks averaging over 40 seconds reach 81.7% average success. Caveat: each context-length result uses a separately trained model at that fixed window.
•
If you are building a video-conditioned robot policy and planning to feed in longer history, the paper is strong evidence that your video foundation’s pretraining objective matters more than you might think. A bidirectional video generator that is only adapted to be causal may not reward longer windows. Worth testing an autoregressive backbone before concluding history “does not help” for your task.
•
The predict-then-act pattern (partially denoise future latents, keep them in latent space as conditioning, then denoise actions) is a cheap way to get video-prediction benefits without paying for pixel decoding. Reasonable to try on top of an existing action head.
•
The deployment tricks are generic enough to borrow. Streaming VAE encoding, within-call KV reuse of text and observation keys, and NVFP4 on video layers while keeping action compute in BF16 are each separable from the modeling contribution.
•
The paper does not establish that one trained Long-WAM dynamically adapts its context length. Each window is a separately trained model, and the authors call test-time adaptive windows a next step rather than a result.
•
Pairing with a high-level planner (they use GPT-6 Astra) lifts composite-task success substantially (31.4% to 54.4% overall on RoboCasa365), but this is a hierarchical setup, not an end-to-end claim. If your tasks are atomic skills, the planner may not matter.
•
Code and project page: GitHub.
The context-scaling comparison trains a separate model per window length, so it measures the ceiling of what each window can do, not the behavior of a single flexible-window policy. The 38.4s regression is almost certainly a data-coverage artifact (80% padding), so the paper does not actually probe where the predictive benefit saturates. Real-world baselines (\pi_{0.5}, Fast-WAM) are evaluated in their native configurations rather than retrained on the same data, so the dramatic 0-of-20 failures on dynamic tasks reflect published checkpoints and should not be read as a controlled head-to-head. And the training-time surrogate for the predicted future (ground-truth frames forward-noised to the inference noise level) is an approximation; the authors flag that the inference-time distribution of future latents is not identical.