Research questionHow can robot VLA policies acquire phase-specific causal reasoning without inference-time overhead?Imitation-trained vision-language-action policies can predict what action to take without representing why that action fits the current manipulation phase. Generating rationales or rolling out future states at every step adds computation that compounds over long-horizon tasks.