On-policy distillation from an RL-trained teacher to a student follows a predictable early phase where held-out accuracy grows linearly in square-root KL from the student’s starting point, and peak student accuracy stops improving once teacher size exceeds student size.
Suppose you’ve spent compute doing Group Relative Policy Optimization (GRPO) on a 1.5B math model and it works well. Now you want the same math skill in your 7B and 14B models from the same family. The obvious move is to repeat RL at every scale, which is expensive and gives you no way to predict the outcome before training finishes.
The alternative studied here is On-Policy Distillation: let the big student generate rollouts, and have the small RL expert score each token to shape the student’s policy. Prior work (notably Gao et al. 2023 on reward-model overoptimization) showed that for standard RL against a learned reward, you can model gold-metric dynamics as a function of KL from the starting policy, with coefficients that scale predictably. The authors ask whether OPD admits the same kind of forecast: given student size, teacher size, and teacher score, can you predict what the distilled student will reach before you run it?
The setup is a controlled grid on Qwen2.5 base models from 0.5B to 14B, all SFT’d the same way, with RL teachers trained on a GSM8K+MATH mix. They run 25 teacher-student pairs covering weak-to-strong, same-base, and strong-to-weak directions.
The key move is choosing the right x-axis for training progress. Instead of optimizer steps, they measure how far the student’s policy has drifted from its SFT starting point in token-mean reverse KL, estimated with the K3 KL Estimator. They then plot gold-score (held-out accuracy) against d = sqrt(KL), the square root of that drift. The square root is motivated by local geometry: along a smooth training path, accuracy changes to first order in the parameter step while KL changes to second order, so accuracy is linear in sqrt(KL) near the start. Formally, G(d) ≈ G(0) + m·d for small d, where m depends on how well the update direction aligns with the gold-score gradient.
Every one of the 25 runs shows this linear Useful-transfer regime early on (R² between 0.93 and 0.99 on the first 30 checkpoints). After that, trajectories diverge into saturation, slow improvement, or outright regression, with no shared functional form.
They then fit joint power laws for two targets: peak remaining error (1 − G_peak) and the initial slope m. The predictors are student parameter count, an effective teacher size (teacher capped at student size, since bigger teachers stop helping past that point), and the teacher’s own gold score.
# conceptual loop for one teacher-student pair
student = sft_init(N_student)
for step in range(max_updates):
prompts = sample(dataset)
rollouts = student.generate(prompts) # on-policy
for tok in rollouts.tokens:
reward = log p_teacher(tok) - log p_student(tok) # Vanilla-OPD
student.update(tok, advantage=reward) # zero-discount
d = sqrt(token_mean_kl(student, sft_init))
log(d, gold_score(student))
# fit G(d) = c + m*d on first 30 checkpoints
They also compare Vanilla-OPD against Delta-OPD, which uses the log-ratio between the RL teacher and its pre-RL base as the token reward (isolating what RL added) plus an explicit KL-to-reference penalty.
•
Weak-to-strong transfer works, and the student overtakes its teacher. In every weak-to-strong pair, the student’s peak gold score exceeded the teacher’s own. The 14B student taught by the 0.5B RL expert still gained meaningfully over its SFT baseline, though not as much as under a larger teacher.
•
Teacher size helps only up to student size. For 7B and 14B students, peak accuracy rose monotonically with teacher scale. For the 0.5B student, the 3B teacher was best (40.8%) and the 7B and 14B teachers were worse (37.7–39.0%). The authors cap teacher size at student size in their scaling law to model this.
•
At matched teacher score, the smaller teacher transfers better. They test this directly: a 3B teacher checkpoint caught early in its RL run scores 66.0, slightly above the fully-trained 1.5B teacher’s 63.6. When both teach the 7B student, the 1.5B teacher’s student peaks at 77.5 while the early-3B’s student peaks at 73.8. Only the joint law (conditioning on both scale and score) predicts this ordering; scale-only and score-only laws both wrongly favor the 3B teacher.
•
The peak law extrapolates. Withholding the largest student or teacher scale from the fit, the joint law predicts the held-out peaks within roughly 0.7 accuracy points for Vanilla-OPD and 0.4 for Delta-OPD.
•
Delta-OPD beats Vanilla in most shared cells. Larger matched-KL slope in 15 of 17 pairs, larger peak gain in 12, with the advantage concentrated in weak-to-strong pairs.
•
An off-policy SFT cold start hurts weak-to-strong OPD. One epoch of SFT on the 0.5B teacher’s rollouts dragged the 14B student down by 32 points from its initialization, and subsequent OPD did not recover.
•
Bootstrapping through intermediate sizes did not help. Chains like 0.5B → 1.5B → 3B → 7B → 14B peaked below direct OPD from the 0.5B expert at every stage, even though the intermediate teachers scored higher than the 0.5B RL expert. This reverses the bootstrapping gain reported by Burns et al. 2024 in a different setting.
•
Late-stage regression looks like proxy overoptimization. When gold score falls, the teacher-induced token reward keeps rising, which is the fingerprint of optimizing a proxy past the point where it tracks the thing you care about.
•
If you have a model family and want task expertise at every scale, consider running RL once on a small model and distilling via OPD to the larger ones, rather than repeating RL at each scale. The paper shows this works for math reasoning on Qwen2.5 up to 14B. Whether the specific exponents transfer to other families, tasks, or longer rollouts is untested.
•
Pick teachers by competence per parameter, not raw score. The paper’s clearest decision rule: given two candidate teachers at similar gold score, the smaller one will produce a better student. Given two teachers at the same size, the better-scoring one wins. The joint law gives you an actual prediction before you train.
•
Do not do an SFT warmup on teacher rollouts before weak-to-strong OPD. The cold start pinned students near the weak teacher’s score and OPD did not recover it. Start on-policy from the student’s own SFT checkpoint.
•
Skip bootstrapping chains. Distilling from an intermediate OPD-trained teacher does not beat distilling directly from your smallest RL expert, in this setup.
•
Use token-mean reverse KL from the student’s starting policy as your x-axis for monitoring OPD runs. The linear regime in sqrt(KL) is a clean signal; when a run breaks the predictive band of its early linear fit for several consecutive checkpoints, treat that as the end of useful transfer and stop. The paper gives this a concrete rule (95% band, 3 consecutive deviations) that flagged 9 of 25 runs as having left the regime.
•
Treat the specific power-law coefficients as a scale summary rather than a universal predictor, especially for the transfer rate, whose fit is noisier (log-space R² of 0.59 for Vanilla-OPD).
•
One model family (Qwen2.5), one task domain (grade-school and competition math), one random seed per cell, and rollouts capped at 2K tokens. The authors are explicit that transferability to other families, tasks, or longer-reasoning settings is untested.
•
The peak law’s teacher-size exponent is only identifiable because they added bootstrapped-chain teachers whose score sits off the size trend. Within the clean RL-endpoint grid, scale and score are collinear at r = −0.9999 and cannot be separated.
•
The useful-transfer rate law is noisier than the peak law and does not consistently beat a plain affine baseline under cross-validation. Read it as interpretation, not prediction.
•
The paper does not model where the peak occurs in d, only its value, because peak locations are non-monotone across the grid.
•
Gold-score evaluation on a held-out math test set does not establish safety, calibration, or robustness of a distilled model, and distillation can propagate a teacher’s biases and hallucinations into a larger student that would otherwise not have had them.