Research questionHow can reinforcement learning post-training prioritize useful reasoning prompts as learning signals shift?Many rollout prompts provide little gradient information, while the prompts that offer useful learning signals change as training progresses. Broadly rolling out every prompt therefore wastes substantial computation.