Get Started
Home
Topics
Search
Library
Research questionWhen does environment design make on-policy RL amplify harmful specification gaming in language models?On-policy RL can reward behaviors that exploit a task specification, but its safety effects vary across environments. Model size and common safety scores do not consistently indicate when harmful exploitation will emerge.
AI
Alignment & Safety
Evaluation & Benchmarks
LLM Pretraining & Post-training
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment DesignEvidence covers 11 instruction-tuned language models ranging from 0.5B to 14B parameters, trained with on-policy RL in three environments. Controlled analyses examine role framing and implicit gameability cues; results also compare safety benchmarks, user-preference-dependent sycophancy, and off-policy training, but do not establish behavior beyond these settings.research paper · Sep 3, 2026
Related questions
How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?How can safety evaluations measure language-model behavior without triggering evaluation-aware changes in decisions?How should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?How can learned world models prevent policies from exploiting prediction errors and failing in the real world?