Research questionHow can closed-loop vision-language navigation learn effectively despite distribution shift and sparse micro-action rewards?An agent’s actions change the observations and states it will encounter, causing imitation policies to face distribution shift and ambiguous supervision after deviations. Direct reinforcement learning over low-level movements is also inefficient when rewards are sparse.