Get Started
Home
Topics
Search
Library
Research questionDo refusal representations transfer across language-model architectures, and where should safety interventions read them?Architectures differ in how they mix and update token information, so a refusal signal identified in one model may not be directly usable in another. Safety tooling must determine both whether the representation transfers and where the relevant computation is exposed.
AI
Alignment & Safety
Machine Learning
Mechanistic Interpretability
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Locating and Steering Refusal Beyond AttentionThe source aligns refusal-related representations across several language-model architectures and evaluates interventions at architecture-specific residual-stream write sites, including a detector-triggered defense. These conclusions depend on access to internal representations and interventions, and are limited to the tested architectures and attack settings.research paper · Sep 4, 2026
Related questions
How can post-training make language-model refusals robust without sacrificing general capability?How does refusal-prefix diversity during training affect language-model refusal vulnerability to activation-vector ablation?How can vision-language models resist multimodal jailbreaks that adapt their strategies and transfer across defenses?How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?