Get Started
Home
Topics
Search
Library
6 min read · LLM Training · Reinforcement Learning · Added Sep 30 · Paper published Sep 28, 2026

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Source: research paper via Hugging Face Daily Papers
0:00 / 6:49
Multi-teacher on-policy distillation silently lets the loudest RL specialist dominate: in a Qwen3 4B setup, instruction-following log-ratios spread 2.3-4.4x wider than math, supplying 94% of the gradient. Rescaling each domain’s advantages by inverse std, clipped [0.25,4], recovers the math teacher’s skill with no extra rollouts.
TL;DR
DN-MOPD trains one multi-skill student from several RL-trained specialist teachers by rescaling each domain’s feedback to have the same spread, so that a noisy instruction-following teacher stops swamping quieter math and code feedback in the shared gradient.
Why It Matters
Suppose you’ve done separate RL post-training runs to get three specialists: one great at math, one at code, one at following instructions. You want to ship ONE model that does all three. The current recipe is MOPD: for each training prompt, the student generates an answer, and the specialist matching that prompt’s domain scores the student’s tokens. The token-level gap between teacher and student log-probabilities becomes the learning signal.
The problem the paper flags: this routing decides which teacher grades a given prompt, but it silently assumes every teacher grades on the same scale. It doesn’t. In the authors’ Qwen3.5 setup, instruction-following log-ratios are 2.3 to 4.4x as spread out as the pooled signal, while math log-ratios are about half as spread. Because all three teachers push updates into the same shared parameters, the loud teacher wins. For the initial 4B student, the instruction-following loss supplies 94% of the combined gradient under equal weights. The math specialist’s skill barely transfers to the student, and MOPD ends up no better than a student taught by just one well-chosen specialist.
How It Works
The intuition: before you sum up feedback from three teachers, put them on a common scale so the loudest one doesn’t dominate. DN-MOPD does exactly that with one added step per batch.
For each domain, the paper measures the standard deviation of the teacher-vs-student token log-ratios in the current batch. It also computes a pooled standard deviation across all domains. Each domain gets a multiplier: pooled spread divided by its own spread, clipped to [0.25, 4]. That multiplier scales the domain’s per-token advantages before the standard clipped policy-gradient loss (essentially Group Relative Policy Optimization (GRPO)-style, following the Uni-OPD implementation).
Domains with tight feedback get amplified; domains with dispersed feedback get muted. Signs are preserved, so a token the teacher wanted to encourage still gets encouraged, just weighted differently. MOPD is the special case where every multiplier equals one.
for batch in training: responses = student.sample(batch.prompts) r = teacher_logp(responses) - student_rollout_logp(responses) sigma_all = std(r) # pooled across domains for d in domains: sigma_d = std(r[domain == d]) w[d] = clip(sigma_all / sigma_d, 0.25, 4) A = teacher_logp - student_actor_logp # recomputed advantages A_tilde = stopgrad(w[domain] * A) student.update(clipped_opd_loss(A_tilde))
No extra teacher calls, no learned router, no per-model-size hyperparameter search. The multipliers come from batch statistics you already have from rollout scoring.
What They Found
Across three Qwen3.5 sizes (9B, 4B, 2B), DN-MOPD improves the six-task average over label-routed MOPD at every size and both evaluation lengths, with +1.17 to +2.36 pp at the 16K cap and +2.47 to +3.08 pp at 8K. Paired confidence intervals sit above zero, and the gain holds across three student seeds.
The six evaluation tasks are two math (AIME 25/26), two code (LiveCodeBench v5/v6), and two instruction-following (IFEval, IFBench). Math is where MOPD fails hardest and where DN-MOPD helps most. At the 16K cap, plain MOPD shows no math gain over the initial student at any size, while DN-MOPD recovers most of what the math specialist could offer.
The most informative ablation: fixed weights. Amplifying math alone (2x, others 1) recovers only about half of DN-MOPD’s gain at 4B and little at 2B. Reducing instruction-following alone (0.25x, others 1) recovers most of it. So the mechanism is mainly turning down the noisy teacher, not turning up the quiet one. Also: fixed weights close to what DN-MOPD measures on the first batch work about as well as per-batch estimation at 9B and 4B, so per-batch calibration is more of a convenience than a hard requirement at larger sizes.
DN-MOPD does NOT beat every baseline. Offline SeqKD on fixed teacher answers, and simple parameter-space task arithmetic, both reach higher totals in absolute terms. The paper’s claim is narrower: given that you’re doing MOPD, this fixes a real defect in it. The authors also note a boundary case: on an earlier Qwen3-4B setup, DN-MOPD showed no clear gain, and its measured math multiplier there was 1.0, meaning that construction didn’t have the imbalance to fix.
What’s Useful
If you’re running multi-teacher on-policy distillation and combining specialists trained by different RL pipelines, check the per-domain standard deviation of teacher-minus-student log-ratios on your first training batch. If one domain is several times more dispersed than the others, your shared update is probably being driven by that domain regardless of how you balanced prompt counts. This is the concrete failure DN-MOPD targets.
A cheap intervention worth testing before implementing DN-MOPD: pick fixed per-domain loss weights inversely proportional to the measured first-batch spreads, clipped to something like [0.25, 4]. The paper’s controls show this matches DN-MOPD at the two larger sizes. Per-batch re-estimation added value only at the smallest model.
If your specialists were all trained by the same recipe (same reward shape, same clipping, similar KL), you may not have this imbalance. The Qwen3-4B boundary result shows the fix produces no gain when there’s nothing to fix. Measure before adding machinery.
Don’t take this as evidence that DN-MOPD beats offline distillation or parameter merging in general. Under the paper’s own compute-mismatched comparison, SeqKD-SFT and task arithmetic reach higher totals. The contribution is specifically to the label-routed MOPD recipe that several frontier labs have adopted, not a claim of best-in-class multi-skill integration.
Code and training records are released at GitHub.
Caveats
One model family, one expert pool per size, one prompt mixture. The three student seeds vary sampling order but not the underlying expert training. Cross-family generalization is untested.
The instruction-following multiplier sits at the 0.25 clipping floor in most batches, so the rule bounds rather than equalizes that domain’s scale. Different clipping bounds could change results and were not tuned per size.
The reported gains are relative to MOPD variants. Against offline distillation on fixed teacher answers, or plain parameter-space task arithmetic, DN-MOPD trails on the six-task total under the compute budgets used. Whether the ordering would change under matched compute isn’t resolved.
The scale statistics depend on domain token composition and response length. Instruction-following contributes only ~1% of response tokens but up to half the pooled variance, so the estimator is sensitive to how domains are mixed and how long answers are.
Topics
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Reinforcement Learning93 episodes
LLM Training134 episodes