Research questionHow can image generation and editing agents reliably verify and integrate retrieved multimodal world knowledge?Parametric knowledge may be incomplete for the facts and visual appearances required by a prompt, while retrieved text and images can be difficult to check and combine during generation or editing. As a result, relevant evidence may still fail to prevent factual or visual errors.