On-policy distillation acts like RL where the teacher is an implicit reward model scoring student rollouts. It reweights behaviors the student already produces, so if the teacher scores pathological rollouts highly, training amplifies them into 99% truncated, 38% repetitive outputs even when the teacher itself rarely generates them.
If you post-train a reasoning model, you probably use one of two recipes. Classical distillation copies the teacher’s own completions. On-Policy Distillation instead lets the student generate, then nudges the student’s token distribution toward the teacher’s at those same prefixes. Big labs have adopted it because the student practices on text it would actually produce, not on teacher text it struggles to imitate.
The catch, reported by prior work and reproduced here: sometimes OPD runs collapse. The student starts writing 8000-token answers that repeat "Final Answer: 116" 365 times before truncating. Loss curves look fine. Nothing in the usual training telemetry warns you. Before this paper, the common guess was that teacher-student distribution mismatch was corrupting the KL target. That diagnosis suggests picking a closer teacher, which is expensive and not always possible.
The core reframing: the token-level reverse-KL divergence loss used in OPD is algebraically equivalent to a policy-gradient update where the teacher’s log-probability of each student-generated token acts as a reward, plus an entropy term. So OPD isn’t really “imitating the teacher.” It’s doing RL with the teacher as a reward model, evaluated on student rollouts.
That reframe has two consequences. First, OPD can only reinforce behaviors the student already samples. It can’t teach genuinely new skills, only make existing ones more likely. Second, the reward signal’s quality depends on the teacher’s judgments on student text, not the teacher’s own generation quality. A teacher that writes clean solutions can still assign high scores to garbage student output, because scoring and generating are different operations.
The authors verify this by looking at which S_INIT (initial student) rollouts a given teacher scores highly, before any training happens. In their good setting (JustRL-1.5B teaching DeepSeek-R1-Distill-Qwen-1.5B), top-decile responses are correct and concise. In their bad setting (Qwen3-4B teaching Qwen3-1.7B-Base), top-decile responses are already overlong and mechanically repetitive. OPD then amplifies whichever pattern the teacher happened to prefer.
In pseudo-code, one OPD iteration looks like:
for prompt in batch:
rollouts = student.sample(prompt, n=8) # student-generated
for y in rollouts:
for t, token in enumerate(y):
# teacher scores the student's token at student's prefix
r = teacher.logprob(token | prompt, y[:t])
adv = r - student.logprob(token | prompt, y[:t])
loss += ppo_clipped_surrogate(adv, student)
loss += kl_coef * kl(student || s_init)
student.step(loss)
Note the teacher never generates. It only grades.
OPD doesn’t expand what the student can solve, it makes already-solvable problems easier to sample. On AIME24-26, pass@1 improves substantially after OPD, but the gain shrinks as sampling budget grows, and pass@256 is roughly equal between initial and trained student. A manual audit of problems that appeared solved only by the OPD student found zero cases that survived scrutiny: either the initial student also solved them with more samples, or the OPD solution was invalid despite a correct final answer. On problems the initial student could already solve, OPD lifts per-problem success rate on ~85-91% of them.
Collapsed runs converge normally. Loss and advantage curves for the collapsed Qwen3-4B run look as clean as the successful JustRL run. Trained-student outputs still sit in the low-NLL region of S_INIT (initial student), confirming the student is exploiting pre-existing modes, not discovering new ones.
The collapse is reward hacking, not optimization failure. The authors define a selection gap measuring how much the trained student’s probabilities shift toward teacher-preferred vs. teacher-dispreferred S_INIT (initial student) responses. In JustRL the gap grows modestly (near 0 to 0.07). In the Qwen3-4B collapse the gap balloons to ~26. The student is faithfully learning the teacher’s preference ranking. The preference ranking is just bad.
Two mitigations work, keeping the teacher fixed. Masking the loss on rollouts that hit the generation length cap recovers from collapse mid-training and improves four-benchmark average accuracy by +3.08 pp with the Qwen3-4B teacher. Initializing from an SFT checkpoint (released by prior work, not trained here) prevents collapse from starting and reaches the best average, 16.38% vs. 10.16% for the base-initialized OPD baseline. Smaller gains hold with Qwen3-8B and Qwen3-30B-A3B teachers, though per-benchmark effects are mixed.
•
If you’re running OPD and seeing length blow up or repetition, the diagnosis to run first is not “my teacher is too strong.” Rank a batch of your current student’s rollouts by teacher advantage and look at the top decile. If those rollouts are already pathological before training, you have a reward-hacking problem, and switching to a closer teacher won’t fix the preference bias.
•
Masking the loss on truncated rollouts is cheap. In the paper’s Qwen3-4B run, this leaves only about 10% of sampled responses with nonzero loss and still improves accuracy. Worth testing as a first intervention when you suspect length exploitation.
•
SFT warmup before OPD is the stronger lever when available, because it changes the rollout distribution the teacher gets to score. The paper used someone else’s SFT checkpoint, so this isn’t a controlled test of “how much SFT,” but the direction is clear: a cleaner initial sampling distribution gives OPD less pathology to amplify.
•
Treat pass@1 gains from OPD as sampling-efficiency wins, not capability wins. If your deployment already does best-of-k or self-consistency with large k, the OPD benefit may be smaller than single-sample benchmarks suggest. Worth measuring pass@k on your own task before committing to the training cost.
•
Teacher generation quality is not a reliable proxy for teacher scoring quality on student text. Evaluate a candidate teacher by having it rank your student’s rollouts and checking whether its preferences correlate with correctness, not by looking at how well the teacher solves the task itself.
Code: github.com/HancCui/opd_hacking.
•
All experiments are on competition mathematics with 1.5B-4B parameter models. Whether the reward-hacking framing and the specific mitigations transfer to larger models or non-math domains is open.
•
The coverage audit is deliberately one-sided: extra sampling was spent only on problems where the OPD student appeared to solve something the initial student didn’t. This can only shrink the gap in the direction the paper’s claim predicts, so “no OPD-only problems survived” is an upper bound on initial-student coverage, not a symmetric comparison.
•
The warmup intervention uses an externally released SFT checkpoint, so the paper isn’t claiming a specific SFT recipe. It only shows that a better starting distribution helps.
•
Mitigation gains are not uniform. With the Qwen3-8B teacher, masking lifts the average by only 0.51 pp and lowers AIME25 and AIME26 accuracy. The intervention is targeted at the specific truncation-plus-repetition failure mode observed here, not a general-purpose OPD fix.