Research questionHow can text-to-image diffusion models use multiple candidates and continuous rewards beyond pairwise preferences?When several images are generated for one prompt, reducing their graded reward scores to a winner-loser comparison discards information about their relative quality. That loss can limit how effectively the diffusion model learns the intended preferences.