Research questionHow can reinforcement learning post-training for diffusion models avoid objective mismatch and high-variance updates?Diffusion models are commonly pretrained with score- or flow-matching objectives, while some reinforcement-learning post-training methods optimize a different objective. This mismatch can produce noisy estimators, increase variance, and slow convergence.