Research questionHow can image-text retrieval focus on caption-described attributes while ignoring unmentioned visual information?Image embeddings can preserve visual attributes that captions do not mention, causing similarity to reflect content irrelevant to the retrieval query. This information imbalance can weaken alignment between image and text representations.