RIDE distills a reinforcement-learning-trained teacher into a student by pushing the student’s hidden states past the teacher along the direction RL moved the teacher from its pre-RL base, giving a deterministic per-layer gradient instead of a noisy sampled-token signal.
Suppose you’ve spent a lot of compute RL-tuning a reasoning model, and now you want to fold that improvement back into your base checkpoint (or merge several RL-tuned experts that share a base). The standard move is On-Policy Distillation: let the student generate rollouts, then match the teacher’s next-token distribution on those rollouts.
A recent variant goes further. Treat the teacher’s log-probability advantage over the base as an implicit reward, and up-weight it so the student can surpass the teacher. Call that Extrapolated On-Policy Distillation. In practice this is unstable: crank the extrapolation coefficient up and the student’s outputs lengthen, format breaks, and accuracy drops below the plain teacher-matching baseline. The authors show two concrete reasons this happens, and both point at the same culprit: output space is a lossy, noisy place to measure what RL actually changed.
The practical stakes: if you want “teacher plus a bit more,” you need a signal that (a) sees the full RL-induced change, not what survives the final Language-model head, and (b) doesn’t get noisier as you push harder.
Start with the picture. Three models share the same architecture and tokenizer: the pre-RL base, the RL-trained teacher, and the student (initialized from base). Run all three on the same student-generated prefix. At every transformer layer and every token position, compute the vector h_teacher - h_base. That vector is what the authors call the RL-induced residual: it isolates what RL changed about the model’s internal computation on this specific context.
Standard On-Policy Reverse Distillation (OPRD) trains the student’s hidden states to match the teacher’s. RIDE changes one thing: instead of aiming at the teacher, aim at a point past the teacher, along the residual. One scalar \u03bb controls how far past. At \u03bb=1 you get OPRD back. At \u03bb>1 the student is pulled beyond the teacher in the same direction RL was already pulling.
for prefix in student_rollouts:
h_T = teacher.hidden_states(prefix) # frozen, all layers
h_B = base.hidden_states(prefix) # frozen, all layers
delta = h_T - h_B # RL-induced residual
target = h_T + (lam - 1) * delta # extrapolated target
h_S = student.hidden_states(prefix)
loss = mean_squared_error(h_S, stop_gradient(target))
loss.backward()
Two properties fall out. First, viewed as optimization at a single hidden state, this is equivalent to maximizing a linear reward (how aligned the student’s displacement from the teacher is with the residual) under a quadratic penalty that keeps the student near the teacher. It’s the hidden-state analog of KL-constrained reward maximization, but with the KL replaced by squared distance and the log-ratio reward replaced by an inner product.
Second, and this is the key argument against doing extrapolation in output space: the Language-model head is anisotropic. Its weakest singular directions carry about 79.8% of the residual’s hidden-state energy but only 30.0% of its logit energy. So an output-space loss supervises most of the RL-induced change at a small fraction of its true weight, and sees nothing at all in layers below the final one.
Third property, about noise. In output space, extrapolation uses a sampled-token advantage whose variance scales with (\u03bb-1)^2 and does not vanish as the student approaches the teacher. In hidden-state space the residual is computed from two forward passes, not sampled, so changing \u03bb changes where you’re aiming but not how noisy the aim is.
Four base/teacher pairs: DeepSeek-R1-Distill-Qwen-1.5B with its JustRL counterpart, plus Qwen3-4B, Llama-3.2-3B, and Phi-4-mini each paired with an RL-tuned version trained by the same JustRL recipe. Evaluation is Avg@16 on AIME 2024, AIME 2025, and AIMO (AMC 2022\u20132023), with 3 seeds.
•
RIDE is the only method whose mean sits at or above the RL-trained teacher on all four pairs. On three of the four, the margin is within one across-seed standard deviation, so “approaches or exceeds” is the honest framing.
•
RIDE beats OPRD by 0.97 to 4.06 points across the four pairs. Since RIDE and OPRD differ only in the target (teacher vs. teacher-plus-residual), this isolates the benefit of the residual extrapolation itself.
•
ExOPD, the output-space version of the same idea with the same \u03bb, falls below its teacher on every pair. On Qwen3-4B it falls 14.1 points below the untouched student. The damage is worst when the teacher is close to its base (small residual), which matches the variance prediction: noise scales with (\u03bb-1)^2 regardless of how much real signal the log-ratio carries.
•
Coefficient sweep on R1-Distill: RIDE improves over OPRD for every \u03bb in [1.15, 1.35], peaking at \u03bb=1.25. Beyond \u03bb=1.35 it degrades gracefully. ExOPD at the same \u03bb=1.25 already drops below its own teacher-matching baseline.
•
Direction ablation. Replacing the residual with a random vector of equal norm, a reversed displacement, a residual computed against an unrelated model (Qwen2.5-Math-1.5B-Instruct), or a residual computed on mismatched trajectories all land within one point of OPRD (53.90 to 55.12). RIDE with the real RL-induced residual reaches 56.38. The direction, not the magnitude or the mere presence of a displaced target, is what matters.
•
Mechanistic check. The student’s actual update aligns with the residual at cosine 0.954 and overshoots the specified target (projection 1.69 vs. intended 1.25), so \u03bb sets a direction and target more than a hard destination.
•
If you’re distilling an RL-tuned model back into its base checkpoint and you have both checkpoints, RIDE is a drop-in modification of hidden-state distillation: one extra frozen forward pass of the base, one scalar hyperparameter, no change to the training loop’s communication pattern. The authors recommend \u03bb=1.25 with loss rescaling by \u03bb^-2 so the initial loss matches OPRD. GitHub.
•
If you’ve been using output-space reward extrapolation and seeing instability at coefficients much above 1, the sampled-advantage variance analysis explains why, and predicts the collapse will be worst when your teacher is close to its base. Switching to a hidden-state formulation (if you have the base checkpoint and can read internal activations) is worth testing on your own task.
•
If you can’t access hidden states (hosted-API teacher, for example), none of this applies. RIDE requires that student, teacher, and base share architecture, tokenizer, and head, and that you can run forward passes of all three on identical prefixes.
•
Worth testing, not established: whether the gain holds outside mathematical reasoning, under RL recipes other than JustRL, or when the teacher’s head drifts more than the ~1.65% observed here. The paper evaluates only math competition problems with one RL recipe.
•
Three of four pair-level wins over the teacher are within one across-seed standard deviation. “Approaches or exceeds” is the right reading, not “beats.”
•
The Llama-3.2-3B results sit near the evaluation floor on AIME24 and AIME25 (teacher gets 0.42% on AIME25), so most of the signal on that pair comes from AIMO.
•
The student is trained toward hidden-state targets no model actually produced. The authors flag that calibration and safety properties of such a student are not guaranteed to match the teacher’s and need separate evaluation before deployment.
•
RIDE needs the pre-RL checkpoint. If you only have the RL-tuned model, the method doesn’t apply.
•
A single global \u03bb is used throughout; per-layer or adaptive schedules weren’t explored.