Research questionHow can language models answer sensitive prompts helpfully without compromising safety?Sensitive prompts can contain legitimate informational requests, yet models may respond with outright refusals or generic safety language instead of addressing those needs. The central difficulty is distinguishing safe assistance from responses that would create safety risks.