In strong-to-weak LLM distillation, swapping between on-policy (student-generated) and off-policy (teacher-generated) rollouts barely moves accuracy, forgetting, or update sparsity. KL direction and learning rate explain almost all the variation people have been attributing to rollout policy.
If you’re post-training a small model, you’ve seen the claim that on-policy methods like RL with verifiable rewards beat plain Supervised Fine-Tuning because the model learns from its own outputs. People credit on-policy rollouts specifically with three things: less catastrophic forgetting, sparser weight changes, and better generalization. On-policy training is expensive. You have to generate fresh samples from the current student every step, instead of reusing a fixed dataset of teacher outputs. If on-policy is actually necessary, you pay. If it isn’t, you’re burning compute for no reason.
The problem with the existing evidence: SFT-vs-RLVR comparisons change many things at once (objective, reward signal, supervision density, learning rate), so you can’t tell which knob caused the win. The authors use strong-to-weak Knowledge Distillation as a cleaner testbed, where you can hold everything else fixed and only flip the rollout source.
The setup is standard distillation: a bigger teacher’s next-token distribution supervises a smaller student. Two independent design choices get disentangled. First, rollout policy: who generates the token sequences the student trains on, the student itself (on-policy, OnPD) or the frozen teacher (off-policy, OffPD). Second, KL direction: at each prefix, do you minimize Forward KL or Reverse KL divergence between student and teacher. Conventionally forward KL is paired with teacher rollouts and reverse KL with student rollouts, but there’s no mathematical reason they must be coupled once you compute KL over the full vocabulary. The authors cross all four combinations and sweep learning rate on top.
They also build a rollout-policy spectrum controlled by a parameter the authors call lambda. Think of it as a dial: lambda = -1 is pure teacher sampling, lambda = +1 is pure student sampling, lambda = 0 is a geometric midpoint, and values outside [-1,1] extrapolate toward tokens one model prefers much more than the other. This lets them trace a smooth curve rather than compare two discrete endpoints.
The theoretical core is a gradient analysis. For forward KL, the gradient on student logits is just (student prob) minus (teacher prob), bounded in [-1,1] per coordinate. For reverse KL, the gradient is weighted by the student probability times the log-ratio of student to teacher probability, which can blow up when the student assigns real mass to a token the teacher considers nearly impossible.
# per-token logit gradient at a prefix
def fkl_grad(pi_S, pi_T):
return pi_S - pi_T # bounded in [-1, 1]
def rkl_grad(pi_S, pi_T, D_rkl):
log_ratio = log(pi_S) - log(pi_T) # UNBOUNDED
return pi_S * (log_ratio - D_rkl)
They prove forward KL’s expected update changes at most linearly in the total-variation distance between the two rollout distributions, with a constant that doesn’t depend on how extreme any student/teacher probability ratio is. Reverse KL admits no such bound: an arbitrarily small rollout change can produce an arbitrarily large gradient change.
Experiments distill Llama-3.1-8B into Llama-3.2-1B (and Qwen2.5-7B into Qwen2.5-1.5B) across medical, scientific, and arithmetic reasoning (MedReason, Science, Countdown-3), with three seeds per config.
•
Final accuracy on the trained task: best mean across the three datasets is 72% for OnPD vs 73% for OffPD. Essentially tied. Forward KL stays in 71\u201373% across all rollout and learning-rate choices; reverse KL swings from 35% to 72% depending on learning rate.
•
Catastrophic forgetting (measured on seven held-out benchmarks including MMLU-Pro, IFEval, TruthfulQA, HumanEval): at the low learning rate, mean out-of-distribution score barely moves (within 1.3 pp). At the high learning rate, it drops 11.2\u201314.0 pp. Rollout policy makes only modest differences compared to the learning-rate effect.
•
Update sparsity (fraction of weights that moved by less than 1e-6): 85.3\u201389.6% at the low learning rate, 51.9\u201360.0% at the high one. OffPD updates are at least as sparse as OnPD in every matched comparison. On-policy rollouts do not produce sparser updates.
•
Spectrum sweep (lambda from teacher-only to student-only): forward KL varies by only 5.2 pp of held-out accuracy across the entire spectrum. Reverse KL swings wildly and sometimes collapses, especially at higher learning rates. This matches the gradient-bound theory.
•
Where OnPD actually helps: generalization to a harder variant of the arithmetic task (Countdown-4E, four operands instead of three) is 10\u201315 pp higher pass@k with more on-policy rollouts under both KL directions. Also, OnPD + reverse KL uniquely suppresses an incidental style transfer: when the teacher is instructed to reason in Spanish, that config keeps the student mostly English while the others switch almost entirely.
•
One important caveat on the OnPD generalization win: when they run RL with verifiable rewards on top of the distillation checkpoints, OffPD checkpoints end up with the strongest sustained reward on Countdown-4, even though they started lower. The OnPD head start doesn’t persist.
•
Robustness checks: removing gradient clipping, switching to sampled rather than full-vocabulary KL estimators, and training on longer-rollout math (NuminaMath / MATH-500) all preserve the main story. Reverse KL gets more unstable without clipping, confirming the gradient analysis.
•
If you’re doing strong-to-weak distillation and comparing training recipes, treat OffPD as the baseline. The evidence here says you don’t get accuracy, forgetting, or sparsity for free from going on-policy, and OffPD is substantially cheaper because you can precompute teacher rollouts once. The authors explicitly ask future on-policy distillation papers to include OffPD as a baseline.
•
Pick KL direction based on what you care about, not based on rollout policy. Forward KL is more forgiving: it tolerates a wide range of rollout policies and learning rates, gives higher pass@k (better output diversity), and doesn’t need gradient clipping as badly. Reverse KL is sharper and more fragile but useful if you want to suppress incidental teacher behaviors you don’t want transferred (style, language) and are willing to use on-policy rollouts to stabilize it.
•
Tune learning rate before anything else when you care about forgetting or sparsity. Both are governed almost entirely by learning rate in these experiments. Lower LR with forward KL got top target-task accuracy while barely touching out-of-distribution performance. The paper calls back to a “cliff” effect from prior SFT work: two checkpoints with similar final accuracy can forget very differently.
•
If your real goal is generalization to harder variants, on-policy rollouts are worth testing, but don’t assume the win survives a subsequent RL stage. The paper’s RLVR experiment shows OffPD can catch up and surpass, so measure end-to-end rather than at the distillation checkpoint.
•
Scope limits worth taking seriously: students here are at most 1.5B parameters and rollouts are at most ~2,000 tokens. The longer-rollout math experiment is a single run per condition and hints that OnPD + reverse KL might win in that regime, but the authors label it suggestive, not conclusive.
•
The controlled setting uses small students (1B\u20131.5B) and relatively short reasoning traces. Behavior at larger scale or with much longer chains is not established.
•
Teachers are kept fixed and are from the same model family as students (shared tokenizer). Cross-family teacher/student pairs could behave differently.
•
Full-vocabulary KL is the default. A sampled-KL ablation preserves the conclusions, but production pipelines using heavier sampling-based estimators should still verify.
•
The generalization advantage of on-policy rollouts is demonstrated on one task family (Countdown variants) and does not reliably survive subsequent RLVR. Don’t over-generalize from the pass@k numbers.
•
The style-transfer finding (OnPD + reverse KL preserves student language) uses one judge model and one run per condition. Treat it as a lead, not a settled result.