Get Started
Home
Topics
Search
Library
Research questionWhy does relocating a continuation-triggered suffix bypass aligned LLM safety defenses?Aligned LLMs can respond differently when the same continuation-triggering instruction suffix is moved. This makes the source of these jailbreaks and reliable safeguards difficult to determine.
AI
Alignment & Safety
Inference Optimization
Mechanistic Interpretability
Natural Language Processing
Latest papersRecent research connected to this question, newest first.The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMsApplies to LLM harmful-generation behavior under continuation-triggered jailbreak prompts. The evidence uses attention-head-level mechanistic analysis, causal interventions, activation scaling, and behavioral comparisons across model architectures; it also reports an inference-time steering and distillation strategy with no additional computational overhead.research paper · Sep 4, 2026
Related questions
How can black-box LLMs resist jailbreaks without weight access or retraining while preserving benign-query utility?Can intermediate LLM activations guide faster jailbreak search without weakening attack effectiveness?How can we detect when LLM-decompiled code diverges from original behavior or erases disclosed vulnerabilities?How can LLM trajectory evaluations distinguish genuine prefix value and early outcome information from compute and difficulty confounds?