Get Started
Research questionHow can post-training make language-model refusals robust without sacrificing general capability?Language models may refuse harmful requests for very different internal reasons, and safety behavior concentrated in fragile components may be difficult to trust or modify. Post-training can therefore affect not only refusal rates but also the reliability and controllability of the underlying behavior.
Alignment & Safety
LLM Pretraining & Post-training
Mechanistic Interpretability
Latest papersRecent research connected to this question, newest first.Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering RobustnessThe evidence compares supervised fine-tuning, reasoning-augmented fine-tuning, and ORPO preference optimization across Llama-3.1-8B, Gemma-2-9B, and Qwen3-8B. It examines internal refusal structure, steering robustness, general capability trade-offs, and the possibility of small targeted edits; the results do not establish a method that reliably achieves all desired properties for security-critical deployment.research paper · Sep 3, 2026
Related questions
How does refusal-prefix diversity during training affect language-model refusal vulnerability to activation-vector ablation?Do refusal representations transfer across language-model architectures, and where should safety interventions read them?How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?Can post-training ternarization make language models smaller without unacceptable capability loss or slower inference?
Home
Topics
Search
Library