Research questionCan small targeted grayscale patches force chosen semantics in infrared vision-language models across tasks?Localized perturbations may cause an infrared multimodal system to produce a selected class, caption, or answer instead of reflecting its input. The extent to which this vulnerability transfers across tasks and model architectures is unclear.