Research questionHow should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?Binary safe-or-unsafe metrics can conceal meaningful differences in how models respond as harmful requests become less explicit or more adversarial. Models with similar attack success rates may therefore exhibit substantially different safety behaviors.