Get Started
Home
Topics
Search
Library
Research questionHow can multimodal models detect harmful intent when benign text is paired with risky visual content?A multimodal model may respond safely to a text-only request yet miss harmful intent when the same benign wording is grounded in a risky image. Safety behavior can therefore depend on whether dangerous intent is expressed through language or vision.
AI
Alignment & Safety
Mechanistic Interpretability
Multimodal Models
Latest papersRecent research connected to this question, newest first.Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language ModelsThe source studies multimodal large language models through unsafe-response analyses, representation and attention analysis, and experiments across benchmarks and models. It reports a lightweight representation-transfer intervention using a frozen MLLM backbone, but the supplied evidence does not specify deployment conditions or the full range of visual inputs.research paper · Sep 3, 2026
Related questions
How can multimodal models resist harmful intent unfolding across multi-image, multi-turn conversations?How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?How should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?How can multimodal models rely on images or audio rather than language shortcuts?