Research questionHow can video generators align with human preferences while avoiding costly RL and post-distillation model collapse?Human-preference alignment for video generation often relies on computationally expensive reinforcement learning. Combining it with distillation is difficult because applying reinforcement learning first is costly, while applying it afterward can destabilize or collapse the distilled model.