Get Started
Research questionHow does refusal-prefix diversity during training affect language-model refusal vulnerability to activation-vector ablation?Refusal behavior can be encoded in a concentrated activation direction or low-dimensional subspace, allowing ablation to suppress refusals while largely preserving other capabilities. Repetitive refusal starts may concentrate the gradients and features responsible for this behavior, whereas varied starts may produce a different geometry.
Alignment & Safety
LLM Pretraining & Post-training
Mechanistic Interpretability
Latest papersRecent research connected to this question, newest first.Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacksEvidence comes from frozen-model analyses and controlled synthetic fine-tuning, centered on OLMo-2-0425-1B-Instruct. The analysis examines refusal directions and subspaces, gradient and activation-change stable ranks, and vulnerability to vector-ablation attacks; it does not establish that the findings generalize to all models or attack types.research paper · Sep 3, 2026
Related questions
How can post-training make language-model refusals robust without sacrificing general capability?Do refusal representations transfer across language-model architectures, and where should safety interventions read them?Which internal circuits do steering vectors alter to induce refusal in large language models?How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?
Home
Topics
Search
Library