Research questionCan automated alignment research mitigate multiple measurable safety failures without sacrificing general model capability?Alignment failures such as deception, sycophancy, and jailbreaks can be measured, but reducing several simultaneously may interfere with a model’s broader capabilities. It is also unclear whether automated researchers can develop effective interventions without extensive human guidance.