Research questionShould reasoning LLMs combine on-policy distillation and verifiable-reward RL jointly or sequentially during post-training?On-policy distillation provides dense token-level guidance, whereas reinforcement learning with verifiable rewards supplies sparse feedback on solution correctness. Combining these signals in one optimization step may affect solution coverage and refinement differently from staging them.