Get Started
Home
Topics
Search
Library
6 min read · Agents · LLM Training · Added Oct 9 · Paper published Oct 8, 2026

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Source: research paper via Hugging Face Daily Papers
0:00 / 6:44
Training agents with GRPO wastes most rollouts: when all 8 attempts in a group fail (98% of groups on a 2B code model), the gradient is zero and the batch teaches nothing. Self-Retrospection Distillation turns failed rollouts into “pitfall” foresight targets via a stop-gradient teacher, lifting that 0% baseline to 60.6%.
TL;DR
Self-Retrospection Distillation (SRD) turns every completed agent rollout into training signal by distilling after-the-fact knowledge of what the task required into the model’s own before-acting prediction, keeping learning alive when all rollouts in a group get the same reward.
Why It Matters
You’re training an LLM agent with RL with verifiable rewards: it tries a task eight times, a verifier scores each attempt, and the method pushes the policy toward attempts that scored higher than their siblings. The dominant recipe for this is Group Relative Policy Optimization (GRPO), which compares rollouts within a group.
There’s a catch the paper quantifies carefully. When all eight attempts fail (or all eight succeed), the within-group comparison is flat and the gradient is zero. Standard practice throws the group away and resamples. On a 2B model doing code, the authors measure that 98% of groups are all-failure, so training barely moves. Even at larger scales, 37-98% of groups are reward-uniform depending on model size. All that interaction, all those tool calls, teach the agent nothing.
A second line of work, on-policy self-distillation (On-Policy Self-Distillation is the baseline here), tries to extract more signal: a stop-gradient teacher that sees a correct trajectory supervises the student that doesn’t. But the target is still the next action. The paper asks a different question: what if hindsight supervised what the agent could have known before starting, not what to do next?
How It Works
The mechanism is a second training target bolted onto whatever base objective you’re using (GRPO, OPSD, or their hybrid).
For each task, the model produces two short predictions from the problem statement alone, before any tool calls: a Knowledge block (what might this task require?) and a Pitfall block (how might attempts fail?). These are the foresight. After the rollouts finish, a stop-gradient copy of the same model, now shown the completed trajectory plus the gold answer plus an error annotation, writes a hindsight version of the same block. The training loss pulls the foresight’s token-level distribution toward the hindsight’s.
The clever structural choice: reward decides which channel a rollout feeds. Successful rollouts contribute to the Knowledge target. Failed rollouts contribute to the Pitfall target. So an all-failure group, which GRPO discards, becomes pure Pitfall supervision. An all-success group becomes pure Knowledge supervision. Reward no longer gates whether a trajectory teaches, only what.
The divergence used is generalized Jensen\u2013Shannon Divergence, weighted at 0.01 relative to the base loss. In the main experiments only the Pitfall channel is distilled (Knowledge turns out to add cost without reliable gains, see RQ1 below).
for task in batch: rollouts = policy.sample(task, G=8) # on-policy group for r in rollouts: channel = "knowledge" if r.reward==1 else "pitfall" fore = policy(task, channel) # blind prediction hind = teacher(task, r.trace, r.gold, channel) # privileged loss += jsd(hind, fore) # token-level loss += base_objective(rollouts) # GRPO / OPSD / RLSD
At inference, the foresight is optional. The paper’s main results use it only as a training target.
What They Found
Across 10 benchmarks (math: AIME 2024/2026, AMO-Bench; code: LiveCodeBench, OJBench; search: HotPotQA, 2WikiMultiHopQA, BrowseComp-Plus; agents: ALFWorld, WebShop) and Qwen3.5-4B/9B backbones, adding SRD to GRPO, OPSD, or RLSD (Self-Distilled RLVR) improved category averages in nearly every cell. Highlights, framed as the paper frames them:
•
On the 9B OPSD run, Math average went from 39.70% to 56.92% (+17.22 pp), Search from 47.93% to 58.24% (+10.31 pp).
•
On 4B GRPO, SRD added +10.03, +8.33, and +8.05 pp on Math, Code, and Search averages.
•
The headline reward-uniform result: in the 2B code-only setup with dynamic sampling disabled, baseline GRPO ends at 0.0% success (training peaks at 1.6%); the same budget with SRD reaches 60.6%.
•
Gains persisted when evaluation shifted away from training conditions: training on LCB’s stdin format, testing on functional-format problems and OJBench, SRD still added +8.12 pp on OJBench under 4B GRPO.
•
SRD also stabilized plain OPSD, which at 9B had regressed below the untrained policy on HotpotQA, 2Wiki, and LCB-v6; adding SRD put those columns back above the base model.
Ablations clarify the mechanism. RQ1 shows the Pitfall channel’s distillation loss falls 30-47% over training (it’s learnable) while Knowledge’s stays flat (redundant and slower); Pitfall-only is the recommended default. RQ3 shows GRPO’s wasted rollout fraction is U-shaped in capability: weak policies fail everything, strong policies succeed at everything, both lose contrast. RQ4 measures that SRD’s update is positively aligned with GRPO’s own direction (cos +0.81) and extends it 30% further, while on OPSD it acts like an anchor that shortens the step.
The authors are clear that these are observations about trained systems, not isolated causal claims about individual components.
What’s Useful
•
If you’re training agents with GRPO and your rollout logs show lots of all-fail or all-succeed groups, SRD-style auxiliary distillation is worth testing. The published recipe: one extra blind prediction per task, one hindsight prediction from a stop-gradient teacher that sees the trace and gold, JSD between them at weight 0.01. It composes with your existing objective rather than replacing it.
•
The method requires a verifier for each task (same prerequisite as RLVR) and a gold answer or correct trace to condition the teacher on. If you only have outcome rewards without reference solutions, the hindsight channel loses most of its content.
•
Default to the Pitfall channel only. Table 2 shows Pitfall+Knowledge is the one setting that sometimes lands below baseline, and the extra rollouts cost real wall-clock (rollout is 78-87% of a step).
•
Generating the foresight at inference mostly did nothing in Table 3, and in one checked case the predicted foresight was wrong in a way that misled the whole episode. Treat the foresight as a training-time target, not a deployed component.
•
The strongest evidence is in the reward-sparse regime (small models, hard tasks). If your base policy already saturates its benchmark, the headroom shrinks and the per-domain tradeoff visible in Table 3 (ALFWorld task types swing from -20 to +40 pp between arms) becomes more concerning.
Caveats
•
Agentic results trade across task types: no SRD arm improved all six ALFWorld task categories over its base. The average hides this.
•
The paper trains Qwen3.5-4B and 9B only, with a single seed per setting. The authors flag the RLSD-vs-Knowledge interaction as a conjecture, not a measured effect.
•
The 29.9% “hindsight surfaces an interface fact the foresight missed” number is from a lexical pattern match, not blind human annotation; it measures that the interface is mentioned, not that the advice is correct.
•
SRD needs the teacher to be the same model with privileged context. If you only have hosted API access to the model you’re deploying, you can’t instantiate the stop-gradient self-teacher this method requires.
•
The paper’s displacement analysis (RQ4) measures endpoints of training runs that share hyperparameters but not total loss, so it conflates step size with number of effective steps. The authors note this.
Topics
Agents
LLM Training
Reinforcement Learning
Agents
LLM Training
Reinforcement Learning
Up next in Agents
RunningTab: Direct Workspace Interaction with Environment-Side Tabs
nanoMuse: An Open-Source Personal Agent for Every Device You Own
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Agents244 episodes
Reinforcement Learning125 episodes
LLM Training169 episodes