AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2609.32577-gagar.md

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

paperupdated 2026-09-30created 2026-09-30

TL;DR

Binary test-pass rewards make GRPO blind to code quality: every trajectory that passes the tests gets the same advantage, so the policy learns nothing about preferring a clean, scoped fix over one that also rewrites four unrelated files. GAGAR puts all trajectories from a group in a shared workspace, has an SFT-trained agentic grader rank the passing ones, downweights the lower-ranked ones and rescales the rest to preserve the group's original advantage sum. Evaluated at industrial scale on pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B) and MiMo-V2.6-Pro (1.02T): better code-agent performance, reduced trajectory-length growth, and more stable training (source).

Authors & Org

Not stated in the snapshot — no author block, and arxiv.org is blocked from this sandbox. The affiliation is nonetheless legible from the experiments: the paper runs on "pre-RL SFT checkpoints" of MiMo-V2.6-Flash and MiMo-V2.6-Pro, which are Xiaomi's models, and pre-RL checkpoints of a 1.02T model are not something an outside group has. Recorded as inferred, not stated — this page does not name an author.

Method

The diagnosis is exact and it is a property of GRPO rather than of code:

Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements.

So the policy gets no signal favouring clean, targeted implementations over those containing unnecessary or out-of-scope changes. Every passing trajectory is equally right.

GAGAR's pipeline:

  1. Dynamic sampling retains only groups containing both passing and failing trajectories — groups with no contrast carry no learning signal.
  2. All trajectories from a group go into one shared workspace.
  3. An SFT-trained agentic grader inspects them jointly and ranks the test-passing candidates. Joint inspection is the design choice: relative quality is easier to judge side by side than absolutely.
  4. Lower-ranked trajectories are downweighted.
  5. All passing trajectories' advantages are proportionally rescaled to restore the original sum.

Step 5 is the part worth naming. Sum-preserving redistribution means the group contributes the same total gradient magnitude as before; only its internal distribution changes. That keeps the relative weighting from quality-based downweighting while shifting credit toward better implementations, without changing how much the group counts overall — which is why it can be dropped into an existing GRPO setup rather than requiring the loss to be retuned.

Results

  • Controlled code-only Flash experiments: improved code-agent performance, reduced trajectory-length growth, and more stable training.
  • Applied in large-scale mixed-task RL with both Flash and Pro.

No number appears anywhere in this entry. The snapshot's abstract states three directional improvements and quantifies none of them. That is recorded as the state of the capture, and it is the reason this page makes no benchmark claim: an industrial-scale RL result with no figure is a report, not a measurement.

Trajectory-length growth is the most interesting of the three even unnumbered. Length inflation under RL is the code-agent equivalent of reward hacking — an agent that touches more files passes more tests by accident — and it is exactly what a quality-ranked reward should suppress. The paper claims it does.

Significance

It is the reward-side counterpart to a problem this wiki keeps recording on the harness side. Eval Harness Configuration holds case after case of a benchmark number that measures the instrument rather than the model. GAGAR is about the same gap inside training: a binary verifier is a harness, and it scores "passed" identically for a clean patch and a sprawling one. Test-pass RL has been treated in this wiki as the well-behaved case — the one place a real verifier exists — and this paper is the argument that the verifier is coarser than it looks.

Read with Relic: From Multi-Agent Collaboration to Persistent Organizational Capability, also captured today: both target the gap between passing and being a good contribution, one by fixing the reward and one by binding organizational protocol to the runtime. Read with RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling from two days earlier, the pattern is clearer still — three papers in three days manufacturing gradations where a verifier gives only a boolean.

It is also the first artefact this wiki holds from inside Xiaomi's RL stack. Xiaomi was created 2026-09-23 with Key People and Strategic Position deliberately empty, and the MiMo release was noted as open-sourcing its RL stack and training environments. This is a paper describing what that stack does, which is materially better evidence than the release notes.

Open Questions

  • Every number. Not one figure is quoted in the snapshot. "Improved", "reduced" and "more stable" are the whole of the results.
  • Who grades the grader? The agentic grader is SFT-trained, so quality is defined by whatever it was trained on, and that data is not described. The reward has been made richer and also more learned.
  • Does sum-preservation matter? It is presented as the mechanism, but no ablation against non-sum-preserving redistribution is reported.
  • Authors, above — inferred from the checkpoints, not stated.
  • Does 310B or 309B describe MiMo-V2.6-Flash? This paper says 310B total; this wiki holds 309B from third-party reporting. Recorded on MiMo-V2.6-Pro under ## Conflicting Reports rather than silently reconciled.

Cite

arXiv 2609.32577 — Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL, 2026-09-26. HuggingFace Daily Papers, 2026-09-30, 69 upvotes — a popularity signal from that community and not a quality or importance ranking (source).

Referenced by

Sources