Research questionHow can post-training make language-model refusals robust without sacrificing general capability?Language models may refuse harmful requests for very different internal reasons, and safety behavior concentrated in fragile components may be difficult to trust or modify. Post-training can therefore affect not only refusal rates but also the reliability and controllability of the underlying behavior.