Get Started
Home
Topics
Search
Library
Research questionHow can multimodal models resist harmful intent unfolding across multi-image, multi-turn conversations?A harmful objective can be assembled incrementally across turns and images, leaving each individual exchange seemingly benign. Alignment methods designed for isolated interactions may therefore miss risks that become apparent only from the dialogue history.
AI
Alignment & Safety
LLM Pretraining & Post-training
Multimodal Models
Latest papersRecent research connected to this question, newest first.Towards Multi-modal Multi-turn Safety: From Agentic Interaction to Strategic AlignmentThe work studies visual multimodal language models using 11,270 multi-image dialogues and 500 refusal VQA pairs. Reported experiments use Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B on multi-modal multi-turn safety benchmarks, so the evidence is limited to these models, data, and evaluation settings.research paper · Sep 3, 2026
Related questions
How can multimodal models detect harmful intent when benign text is paired with risky visual content?How can vision-language models resist multimodal jailbreaks that adapt their strategies and transfer across defenses?How can multimodal models rely on images or audio rather than language shortcuts?How can speech language models consistently use paralinguistic cues in open-ended, multi-turn dialogue?