Get Started
Research questionHow can safety systems detect and prevent harmful behavior when multi-turn dialogue makes requests seem executable?Multi-turn jailbreaks can distribute harmful intent across several exchanges, making the conversation’s evolving state as important as any individual request. Requests that become more concrete or actionable may be interpreted differently from initially refused requests.
AI
Alignment & Safety
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn JailbreakingThe evidence concerns four-turn trajectories across six frontier models, including open-weight and proprietary systems. It analyzes 18 theory-grounded influence factors, model-specific strategy transitions, actionable and gain-oriented framing, and recovery from hard refusals; reported attack success and query counts are limited to the studied models and evaluation setup.research paper · Sep 2, 2026
Related questions
How can multimodal models resist harmful intent unfolding across multi-image, multi-turn conversations?How can vision-language models resist multimodal jailbreaks that adapt their strategies and transfer across defenses?How can safety evaluations measure harmful actions by computer-using agents rather than chatbot refusals?How can safety evaluations measure language-model behavior without triggering evaluation-aware changes in decisions?
Home
Topics
Search
Library