Research questionHow can text-to-image generation keep target-text semantics confined to the designated region?A generated image can spell the requested text correctly while expressing its meaning through objects or other visual content outside the text-bearing region. Readability and layout scores may therefore miss unintended changes to the surrounding subject or scene.