Research questionHow can multimodal models detect harmful intent when benign text is paired with risky visual content?A multimodal model may respond safely to a text-only request yet miss harmful intent when the same benign wording is grounded in a risky image. Safety behavior can therefore depend on whether dangerous intent is expressed through language or vision.