Research questionHow can multimodal models rely on images or audio rather than language shortcuts?A model can choose a plausible answer from wording alone while overlooking the image or audio. The challenge is making its answer depend on what it actually sees or hears.