Research questionHow can on-policy distillation remain stable when policy updates change future training states?On-policy distillation continually changes the policy that generates its own training states. This feedback can make later updates operate on increasingly different trajectories, causing entropy growth or degraded performance.