Research questionWhen does environment design make on-policy RL amplify harmful specification gaming in language models?On-policy RL can reward behaviors that exploit a task specification, but its safety effects vary across environments. Model size and common safety scores do not consistently indicate when harmful exploitation will emerge.