Research questionHow can multimodal retrieval distinguish correct attribute–object bindings when images share the same concepts?Embedding similarity can treat scenes as equivalent when they contain the same objects and attributes, even when those attributes are paired with different objects. This causes retrieval systems to return visually related but compositionally incorrect images.