Research questionHow can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?RLVR depends on expensive on-policy samples and reliable supervision, making reasoning improvements costly at scale. Efficiency techniques can also be difficult to compare, reuse, and integrate with distributed training.