Get Started
Home
Topics
Search
Library
Research questionCan automated alignment research mitigate multiple measurable safety failures without sacrificing general model capability?Alignment failures such as deception, sycophancy, and jailbreaks can be measured, but reducing several simultaneously may interfere with a model’s broader capabilities. It is also unclear whether automated researchers can develop effective interventions without extensive human guidance.
AI
Alignment & Safety
Evaluation & Benchmarks
LLM Pretraining & Post-training
Machine Learning
Latest papersRecent research connected to this question, newest first.Automated Researchers Can Mitigate Well-characterized Alignment FailuresThe evidence covers automated researchers developing post-training methods across 10 alignment failures, evaluated on targeted and held-out benchmarks, multi-turn behavioral audits, and models up to 4.7 times larger. It also compares these methods with one-shot efforts from 28 experienced researchers and tests whether human ideas improve automated research.research paper · Sep 2, 2026
Related questions
How can post-training make language-model refusals robust without sacrificing general capability?Do newer, larger vision-language models reliably improve autonomous-driving performance without task-specific adaptation?How can robotic manipulation models recognize and recover from failures during task execution?How can we detect and localize failures in long-horizon VLA execution with limited timestamp labels?