Get Started
Research questionHow can black-box LLMs resist jailbreaks without weight access or retraining while preserving benign-query utility?Jailbreak prompts can bypass safety alignment, while defenses that depend on model weights or internal signals cannot be added to many deployed systems. The practical difficulty is reducing malicious compliance without disrupting ordinary user requests.
AI
Alignment & Safety
Natural Language Processing
Latest papersRecent research connected to this question, newest first.AlcaTRAz - Anchored Tree-Rule Defense Against JailbreaksThe described defense transforms input text without modifying or retraining the target model. Evidence covers 33 open-weight models, 22 jailbreak attack types, and short single-turn benign questions; adaptive attackers were not evaluated, and residual high-severity jailbreak failures remain.research paper · Sep 3, 2026
Related questions
Can intermediate LLM activations guide faster jailbreak search without weakening attack effectiveness?Why does relocating a continuation-triggered suffix bypass aligned LLM safety defenses?How can LLM sandbox security remain reliable when linguistic monitoring misrepresents internal computation?How should LLM safety be assessed when jailbreak vulnerability varies by language and persuasive phrasing?
Home
Topics
Search
Library