Get Started
Research questionHow can reinforcement learning post-training prioritize useful reasoning prompts as learning signals shift?Many rollout prompts provide little gradient information, while the prompts that offer useful learning signals change as training progresses. Broadly rolling out every prompt therefore wastes substantial computation.
LLM Pretraining & Post-training
Reasoning
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning ModelThe source studies GRPO-style RL post-training for large reasoning language models across multiple math reasoning benchmarks and models. It uses historical reward trajectories and prompt entropy as selection signals, with evidence focused on rollout efficiency and maintained performance.research paper · Sep 2, 2026
Related questions
How can LLM prompts be automatically refined from recurring reasoning errors without laborious manual engineering?How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?Should reasoning LLMs combine on-policy distillation and verifiable-reward RL jointly or sequentially during post-training?How should open reasoning language models be selected under prompt and deployment-resource trade-offs?
Home
Topics
Search
Library