Research questionHow can visual generative models infer latent rules and produce globally logically consistent images?Visual generators can produce locally plausible details while violating the broader relationships or rules implied by the input. This makes it difficult to determine whether they genuinely reason about visual structure when generating images.