Research questionHow can text-conditioned visual autoregressive models increase image diversity without sacrificing quality?Text-conditioned visual autoregressive models can generate nearly identical images for the same prompt, even when those images are high quality. Increasing variation can sharply reduce image quality, creating a difficult diversity–quality trade-off.