Get Started
Research questionShould reasoning LLMs combine on-policy distillation and verifiable-reward RL jointly or sequentially during post-training?On-policy distillation provides dense token-level guidance, whereas reinforcement learning with verifiable rewards supplies sparse feedback on solution correctness. Combining these signals in one optimization step may affect solution coverage and refinement differently from staging them.
LLM Pretraining & Post-training
Reasoning
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVRThe evidence concerns OPD and RLVR for reasoning language models, evaluated on logic and mathematics benchmarks. It includes analyses of pass@k behavior, learning dynamics, parameter updates, and OPD validation scores as a possible transition signal; broader domains and deployment conditions are not established.research paper · Sep 4, 2026
Related questions
How can on-policy reasoning models use dense self-guidance without reinforcing incorrect solutions?How can reinforcement learning give LLMs useful intermediate feedback when rewards reveal only final correctness?How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?How sparse can token-level supervision be in on-policy post-training without weakening language-model reasoning?
Home
Topics
Search
Library