Research questionHow should causal VLMs preserve access to questions placed before image tokens?In causal VLMs, a question before the image can influence visual processing, yet the answer token may have weak access to that question after reading many image tokens. The resulting image-driven answers can ignore what the model was asked.