Get Started
Home
Topics
Search
Library
Research questionHow can image-text retrieval focus on caption-described attributes while ignoring unmentioned visual information?Image embeddings can preserve visual attributes that captions do not mention, causing similarity to reflect content irrelevant to the retrieval query. This information imbalance can weaken alignment between image and text representations.
AI
Computer Vision
Information Retrieval
Machine Learning
Multimodal Models
Latest papersRecent research connected to this question, newest first.TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language AlignmentThe source studies CLIP models and tests a caption-conditioned representation-editing framework using synthetic captions and natural images. Evidence includes controlled attribute-preservation results and retrieval benchmarks; it does not establish behavior beyond these settings.research paper · Sep 2, 2026
Related questions
How can single-pass image captioning capture fine-grained visual details without multi-stage latency?How can audio-captioning datasets represent fine-grained acoustic detail and perceptual ambiguity for better audio retrieval?How can composed image retrieval distinguish changed, preserved, and removed visual details?How can multimodal retrieval distinguish correct attribute–object bindings when images share the same concepts?