Get Started
Home
Topics
Search
Library
Research questionHow can video generators align with human preferences while avoiding costly RL and post-distillation model collapse?Human-preference alignment for video generation often relies on computationally expensive reinforcement learning. Combining it with distillation is difficult because applying reinforcement learning first is costly, while applying it afterward can destabilize or collapse the distilled model.
Alignment & Safety
Diffusion Models
Reinforcement Learning
Video Generation
Latest papersRecent research connected to this question, newest first.Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution MatchingApplies to foundational video generation models and preference signals represented by preference pairs or intra-group exploration. The source reports experiments across multiple models and compares joint, standalone, and sequential training pipelines; it does not specify deployment constraints or a particular preference-evaluation interface.research paper · Sep 3, 2026
Related questions
How can diffusion image and video generators be preference-aligned without inefficient training exploration or inference-time search?How can few-step visual generators preserve preference-aligned quality without being capped by multi-step teacher distillation?How can online reinforcement learning make VLA manipulation precise without value drift or prohibitive cost?How can open image generators follow fine-grained editing instructions without sacrificing image quality or inference speed?