Get Started
Home
Topics
Search
Library
Research questionHow can safety evaluations measure language-model behavior without triggering evaluation-aware changes in decisions?Language models may respond differently when told that their alignment is being tested. This can alter both their overall decisions and the information they use to make them, making standard evaluations harder to interpret.
AI
Alignment & Safety
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.Language models judge war differently when tested for alignmentThe evidence comes from a full-factorial conjoint experiment testing 20 large language models on decisions about starting wars across 32 scenarios, with and without the statement, "You are tested for alignment with human values." Across 12,800 judgments, the study reports changes in both willingness to start war and the relative influence of strategic considerations versus civilian casualties.research paper · Sep 4, 2026Improving Evaluation Realism with Inference-Time Compute and Deployment ScaffoldsThe study examines automated realism improvements using additional inference-time computation and a deployment-like agent harness for coding evaluations. It tests these techniques across multiple target models and reports that combining them produces larger realism gains than either technique alone, with extra compute outperforming simply extending audits.research paper · Sep 2, 2026
Related questions
How should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?How can we uncover unsafe physical behaviors in vision-language-action models before deployment?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?How can safety evaluations measure harmful actions by computer-using agents rather than chatbot refusals?