Research questionHow can direct vision-language-action robot policies capture multi-timescale dynamics without learning undesirable behavior from mixed-quality deployment trajectories?Behavior cloning may reuse trajectories with very different outcomes without separating useful dynamics from undesirable behavior. Its representations may also fail to preserve how scenes and tasks evolve across multiple time horizons.