Get Started
Home
Topics
Search
Library
Research questionHow can language models answer sensitive prompts helpfully without compromising safety?Sensitive prompts can contain legitimate informational requests, yet models may respond with outright refusals or generic safety language instead of addressing those needs. The central difficulty is distinguishing safe assistance from responses that would create safety risks.
AI
Alignment & Safety
Evaluation & Benchmarks
LLM Pretraining & Post-training
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.SHARD: Safe and Helpful Alignment via Self-Reframing DistillationThis concerns large language models responding to sensitive prompts. The evidence covers the DNA dataset and the English subset of LINGUASAFE across most evaluated model families, with helpfulness and safety assessed in those settings.research paper · Sep 2, 2026SafeMath: Safe Solutions for Unsafe Math Word ProblemsApplies to language models answering arithmetic word problems with embedded harmful or sensitive contexts. The evidence covers audits using ToxicGSM, a dataset of 1.9k such problems, and reported results for existing models and a safety-alignment technique; no broader deployment conditions or model-access assumptions are specified.research paper · Sep 1, 2026
Related questions
How should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?How can language models distinguish trusted instructions from untrusted text to resist prompt injection?How can speech language models protect one user’s private context when responding to another speaker?