Get Started
Research questionHow can we evaluate whether text-to-image models express visual metaphors across domains?Visual metaphors require images to combine elements from distinct domains so that an abstract idea is conveyed. Models that render objects faithfully may still fail to structure the composition or map the concepts metaphorically.
Computer Vision
Evaluation & Benchmarks
Image Generation
Latest papersRecent research connected to this question, newest first.Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image ModelsThe source evaluates 11 representative text-to-image models using 1,500 curated visual metaphors spanning three levels and ten categories, with paired prompts of differing specificity. Its hybrid MLLM-as-judge framework measures metaphorical fidelity through multiple-choice questions and dimension-based perceptual scores; the reported results indicate that even strong proprietary models struggle with compositional structuring and cross-domain mapping.research paper · Sep 2, 2026
Related questions
How can text-to-image models preserve variation in unspecified visual factors under long, semantically dense prompts?Can text-to-image models match web-scale performance using smaller, reproducible datasets and models?How can text-to-SVG evaluation capture semantic errors in ways that align with human judgment?How can visual generative models infer latent rules and produce globally logically consistent images?
Home
Topics
Search
Library