Get Started
Home
Topics
Search
Library
7 min read · Code Generation · Reinforcement Learning · Added Sep 30 · Paper published Sep 26, 2026

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Source: research paper via Hugging Face Daily Papers
0:00 / 9:00
Binary pass/fail rewards in GRPO treat a clean three-line fix identically to a 150-turn hack that weakens assertions. Gagar has an agentic grader rank passing rollouts, then rescales advantages so the group’s total positive credit is preserved—lifting SWE-bench Pro to 62.5% while cutting mean turns from 132 to 112.
TL;DR
Gagar trains code agents with RL by having an Agentic grader rank all test-passing rollouts in a group, then redistributing credit from lower- to higher-quality passes while keeping the group’s total positive advantage fixed.
Why It Matters
When you train a coding agent with RL, the standard recipe is: let the agent edit a repo across many turns, run the project’s tests, and hand back 1 if tests pass, 0 if not. With Group Relative Policy Optimization (GRPO), every passing rollout in a group gets the same positive advantage. Every failing rollout gets the same negative one.
That works for teaching the model to pass tests. It says nothing about how it passed them. Two trajectories can both be green while one makes a three-line fix in the right file and the other rewrites unrelated modules, weakens an assertion, or stumbles around for 150 turns before landing on something that works. Binary rewards treat these as identical, so the policy has no reason to prefer the clean fix. In practice the authors observe this failure directly: their binary-reward baseline on DeepSWE v1.1 peaks then degrades, with trajectory lengths ballooning as training continues.
The question this paper asks is narrow and useful: given a group of rollouts that all pass, how do you push the model toward the good passes without breaking the RL math that makes GRPO stable?
How It Works
The method has two parts: a grader that produces a quality ranking, and an advantage-rewrite rule that turns that ranking into training signal.
The grader. For each task where the rollout group contains both passes and failures (mixed-outcome groups only, following Dynamic sampling), an LLM-based grader gets a shared workspace with the task spec, the repo, every trajectory, every submitted patch, and every test log. It can read code and run checks, not just skim text. It then ranks the passing candidates on five axes: approach suitability, implementation precision, minimality, side effects, and codebase consistency. Ranks get bucketed into three tiers (strong, middling, defective) and mapped to a discount factor f_i ∈ (0, 1], where 1 means “full credit” and lower numbers mean “this pass was ugly.” The grader is an Supervised Fine-Tuning-trained checkpoint of the same MiMo-V2.6-Pro model family being trained, chosen partly because it’s roughly 3× faster than their initial Claude Opus 5 grader (about 600s vs 2000s per group).
The redistribution. The obvious move, multiplying each pass’s advantage by its f_i, has a bug. Passing advantages shrink, but failing advantages don’t, so the group’s total pull becomes net negative. In plain terms: penalties for failing get louder relative to rewards for passing, which the authors show blows up entropy and trajectory length. Their fix is to rescale all passing advantages by a common factor λ so their sum returns to what it was before downweighting:
# per mixed-outcome group A_pass_original = 1 - mean_reward # same for every pass under GRPO S_plus = sum(A_pass_original for i in passing) lam = S_plus / sum(f[i] * A_pass_original for i in passing) for i in passing: A_new[i] = lam * f[i] * A_pass_original for i in failing: A_new[i] = A_original[i] # untouched
Because the same λ multiplies every pass, the ratios f_i / f_j the grader established are preserved, but the total positive credit and the negative credit both stay put. It’s zero-sum inside the passing subset. A safeguard caps λ at 1.5 and re-centers if the cap fires.
What They Found
The controlled study is code-only RL on the 310B-parameter MiMo-V2.6-Flash checkpoint, comparing Gagar against a binary-reward baseline on DeepSWE v1.1 and SWE-bench Pro.
•
Performance and stability. The baseline had to be stopped at step 28: its DeepSWE pass rate collapsed from 58.5% at step 20 to 50.2% at step 28. Gagar hit 62.2% at step 28 and kept climbing to 63.4% at step 44. On SWE-bench Pro, the baseline plateaued around 59% while Gagar reached 62.5% at step 52.
•
Efficiency. At step 28 on DeepSWE, Gagar cut mean interaction turns from 132.3 to 111.6 and mean token length from 191.9k to 172.9k. Similar reductions on SWE-bench Pro. The model is doing more with less thrashing.
•
Quality, judged independently. A blinded rubric evaluation on 30 DeepSWE tasks by Claude Opus 5 (a different model from the training-time grader) gave Gagar a 69.8% average win rate among passing candidates and had it ranked first in 65.0% of groups. Gains concentrated in precision, minimality, and side-effect avoidance, which is what the grader was explicitly asked to reward.
•
The ablation is the important one. Running quality downweighting without the sum-preserving rescale reproduces exactly the instability the theory predicts: policy entropy climbs from 0.358 to 0.905 in 30 steps, mean rollout length from 47k to 114k tokens, and DeepSWE pass rate swings wildly. With redistribution, entropy only drifts from 0.358 to 0.513. This is the cleanest evidence that the mechanism (preserving total positive credit) matters, not just the fact of grading.
•
At scale. Plugging Gagar into large mixed-task RL (coding plus other domains) with Flash and the 1.02T-parameter MiMo-V2.6-Pro yielded avg@3 scores of 67.9 and 71.9 on DeepSWE v1.1 respectively, with Pro coming in above Kimi K3 on DeepSWE and above GPT-5.6 Sol on SWE-bench Pro. Note this is a whole-system comparison against frontier models, not a clean isolation of Gagar’s contribution.
What’s Useful
•
If you’re running GRPO-style RL on any task with a binary verifier and multiple passing rollouts per group, the ablation is the piece to internalize: whatever quality signal you add on top of the verifier, make sure you rescale to preserve the group’s total positive advantage. Otherwise you’re silently amplifying penalties. This is a general point about advantage shaping, not specific to code.
•
If you’re training code agents specifically, the case for a groupwise agentic grader (one that can read the repo and run checks, comparing siblings within a task) is strengthened here relative to per-trajectory scalar reward models. Worth testing whether the same setup helps in your stack; the paper only validates it on MiMo checkpoints with SWE-style benchmarks, so transfer isn’t established.
•
Grading latency is a real cost. Their SFT-trained grader cut it to ~600s per group and they overlap grading with rollouts for other tasks. If you try this with a frontier API grader, budget for the wall-clock hit. The paper doesn’t say what training data was used to distill the grader.
•
Not something you can bolt onto a hosted-API-only setup: the method assumes you own the RL loop, the rollouts, and the grader. The paper does not evaluate a variant that works purely with API access.
•
No code or model release is mentioned in the supplied text.
Caveats
•
The primary comparison is against a plain binary-reward GRPO baseline. There’s no head-to-head against other quality-shaping methods discussed in the related work (PAPO, ReCode, GRRM), so “grading helps” is well-supported but “this particular grading-and-redistribution recipe is best” is not.
•
The rubric evaluation that scores implementation quality uses Claude Opus 5 as judge. It’s a different model from the training-time grader, which reduces circularity, but it’s still an LLM judge with the same five criteria the training grader optimizes. Human-rated quality is not reported.
•
Tier thresholds and the f_i values (0.2, 0.4, 0.85, 0.9, 1.0) are hand-set for the Flash config. No sensitivity analysis is shown.
•
The industrial-scale mixed-task numbers involve many moving parts beyond Gagar (data mix, other domains, base checkpoints). Treat those results as “Gagar is compatible with a large training run,” not as a clean measurement of its contribution at that scale.
•
“Sum-preserving” is exact only when the λ cap doesn’t fire; when it does, the safeguard re-centers and the conservation identities no longer hold strictly.
Topics
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Reinforcement Learning93 episodes
Code Generation44 episodes
NLP92 episodes