Research questionHow can on-policy distillation prevent early student errors from corrupting rollouts in autoregressive vision-language models?In on-policy distillation, the student generates its own token sequence, so an early mistake can alter later states and make subsequent teacher supervision less reliable. This error propagation is especially problematic for autoregressive vision-language models performing structured visual prediction.