Get Started
Home
Topics
Search
Library
Research questionHow can preference optimization align whole-image preferences with token-specific content across spatial and denoising coordinates?Pairwise preferences for complete images do not identify which visual content or denoising stages account for the preference. Applying the same signal broadly can therefore misalign updates with token-specific content.
Diffusion Models
Image & Video Processing
Image Generation
Machine Learning
Latest papersRecent research connected to this question, newest first.ToPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion ModelsThe setting is noise-prediction latent diffusion with attention-based conditioning and image-level pairwise preference labels. Evidence comes from matched three-seed, equal-update U-Net retrainings on SD-1.5 and SDXL, plus an aggregate blind SDXL A/B study; it does not establish an equal-compute comparison.research paper · Sep 3, 2026
Related questions
How can text-to-image diffusion models use multiple candidates and continuous rewards beyond pairwise preferences?How can diffusion image and video generators be preference-aligned without inefficient training exploration or inference-time search?How can attention heads be pruned in text-to-image diffusion transformers without losing prompt-specific object identity?How can preference optimization focus supervision on differing entity slots when contrastive candidates share templates?