Get Started
Home
Topics
Search
Library
Research questionHow can vision-language models resist multimodal jailbreaks that adapt their strategies and transfer across defenses?Static jailbreaks may keep the attack strategy fixed even while changing image or text content. More adaptive attackers can alter both their strategy and parameters, making it harder to determine whether protections generalize across models and defenses.
AI
Alignment & Safety
Diffusion Models
Evaluation & Benchmarks
Image Generation
Multimodal Models
Latest papersRecent research connected to this question, newest first.Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language ModelsThe evidence covers MM-SafetyBench evaluations against GPT-4o, Gemini-3-Pro-Preview, and Seed 2.0, along with transfer to unseen vision-language-model victims and representative defenses. It reports attack success rates but does not establish protection against all models or defense mechanisms.research paper · Sep 3, 2026Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent DebateThe source concerns composite defenses combining text filters, image classifiers, and cross-modal detectors in multiple text-to-image models. Its evidence comes from experiments across multiple models, datasets, and safety configurations, focusing on adaptive jailbreak search under these layered defenses.research paper · Sep 2, 2026
Related questions
How can we assess vision-language models’ instruction following across languages and instruction-hijacking attacks?How can vision-language models maintain visual recognition when modalities are missing and source training data is unavailable?How can vision-language models correct unsafe generations token by token without disrupting safe reasoning?How can deployed vision models resist diverse backdoor triggers without model internals or auxiliary data?