Get Started
Home
Topics
Search
Library
Research questionHow should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?Binary safe-or-unsafe metrics can conceal meaningful differences in how models respond as harmful requests become less explicit or more adversarial. Models with similar attack success rates may therefore exhibit substantially different safety behaviors.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.TIER: Threat Implicitness Benchmark for Evaluating LLM Safety BehaviorsThe source presents TIER, a benchmark spanning four risk domains and four threat levels, from explicit harmful requests to sophisticated jailbreaks. It evaluates six open-weight LLMs using a six-label behavior scale and two independent LLM judges; the reported evidence concerns these models and benchmark settings.research paper · Sep 4, 2026
Related questions
How should LLM safety be assessed when jailbreak vulnerability varies by language and persuasive phrasing?How can language models answer sensitive prompts helpfully without compromising safety?How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?How can safety evaluations measure harmful actions by computer-using agents rather than chatbot refusals?