$ cat wiki/concepts/agentic-rl.md
Agentic Reinforcement Learning
Definition
A paradigm in which an LLM agent learns via RL while interacting with an environment. Instead of a single response, it optimizes sequences of multi-step tool use, planning, and reflection against a reward signal.
Why It Matters
- Squarely at the center of my interests: the intersection of agents + RL + reasoning
- A limitation of plain RLHF (evaluating a single response) — agents require sequence-level evaluation
- One of the core research currents of 2026
State of the Art (2026-10-04)
A controlled study takes the on-policy premise apart, and reassigns each of its three claimed benefits to a different knob. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics varies rollout policy, token-level KL direction and learning rate independently — across Llama3 and Qwen2.5, on scientific, medical and arithmetic reasoning — and reports that rollout policy "does not necessarily play a central role" (source).
Instead: KL direction "more clearly shapes task performance and output coverage", and learning rate "governs forgetting and update sparsity" — the two effects most often credited to on-policy training. The asymmetry is the mechanism: forward KL is "remarkably robust to rollout policy", while reverse KL is "substantially more sensitive and favours student-generated rollouts".
That asymmetry is what makes the result coherent rather than merely contrarian. If the answer to "is on-policy better?" flips with the KL direction, then two honest papers varying both at once can report opposite findings — which is the abstract's stated complaint about prior comparisons: they "vary many factors simultaneously, making the contribution of rollout policy difficult to isolate".
One finding cuts for on-policy and is reported with its own expiry: on-policy data does improve generalisation to harder Countdown arithmetic variants under both KL directions, but "this advantage does not reliably persist after subsequent RLVR". A gain erased by the next training stage is a different claim from one that was never there, and the paper makes the former.
No numeric figure of any kind appears in anything read — no accuracy, no
forgetting rate, no sparsity measure — so every finding above is directional and
no effect size can be assessed. The paper was not read; arxiv.org is blocked
from this pipeline and the HuggingFace Daily Papers abstract is the only text
available. No author or affiliation is stated. Both model families are
open-weight and neither is frontier-class, and the setting is strong-to-weak
distillation only.
State of the Art (2026-09-30)
The verifier is the harness, and a boolean is a coarse one. Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL names the defect precisely: with binary test-pass rewards, GRPO assigns identical advantages to every passing trajectory in a group, so the policy learns nothing that distinguishes a clean, scoped patch from one that also rewrites four unrelated files. This page has treated test-verified code as the well-behaved case — the one domain where a real verifier exists — and this is the argument that the verifier is much blunter than that framing assumes.
GAGAR's answer is sum-preserving advantage redistribution: keep only groups holding both passing and failing trajectories, put the group in one shared workspace, have an SFT-trained agentic grader rank the passing ones jointly, downweight the lower-ranked, then rescale so the group's total advantage is unchanged. Only the internal distribution moves, which is why it drops into an existing GRPO setup without retuning the loss. Reported at industrial scale on pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B) and MiMo-V2.6-Pro (1.02T) — Xiaomi's models — with improved code-agent performance, reduced trajectory-length growth and more stable training, and not one number attached to any of the three (source).
Trajectory-length growth is the item to watch even unnumbered: length inflation under RL is the code-agent form of reward hacking, since an agent that touches more files passes more tests by accident.
Read against RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling below, the two are the same move in two domains — manufacture gradations where the signal is coarse. Video has no verifier at all, so RewardVerse builds a rubric; code has one that returns a boolean, so GAGAR builds a ranker. Both replace a scalar with a structured judgement, and both make the reward more learned in the process.
The infrastructure half, from the same snapshot
Alibaba / Qwen AI Lab's QwenGyre (arXiv 2609.33848, 2026-09-27) is about the cost of running this kind of RL rather than its signal, and its numbers are specific. xlong-horizon tasks are defined as a single execution spanning hours, hundreds of model–environment interactions and nearly 1M tokens per rollout, which breaks online RL two ways: execution variance and rollout delay leave GPUs idle, and non-linear branching floods training with redundant trajectories. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, and its trajectory processor reconstructs branching histories, scores partial progress and deduplicates redundant paths. Scaled to Qwen 3.8 2.4T at 700K tokens per rollout: NL2RepoBench 52.5% → 58.5% (+6.0 points) in 48 steps, and up to 1.85× and 1.78× speedups over Colocate and Async respectively. Recorded as a one-off mention rather than a page, per the rule (source).
The 2.4T figure corroborates this wiki's own. Alibaba / Qwen AI Lab holds Qwen 3.8 Max as a 2.4T MoE from a July capture; the model's own team now states the same number in a paper, which is a stronger source than the reporting it came from.
State of the Art (2026-09-28)
Where there is no verifier, the reward model drifts — and the paper that says so names the failure precisely enough to design against. RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling identifies scalar drift: a video reward model that maps subjective quality straight to one number has its scale collapse or shift across prompts, so the same score means different things in different places and the reward is unusable for RL. The fix is an intermediate representation — generate a dynamic, query-adaptive rubric first, then score against it — trained by RGPO: warm up the scorer on self-evolving seed rubrics, then jointly optimise the rubric generator while realigning the scorer to human ratings. State of the art on the 16-dimensional EvalVerse benchmark, pointwise and pairwise (source).
Why it belongs on this page rather than on a video-generation one. Everything here assumes a reward that holds still. The environments this page tracks get that for free from a verifier — a test that passes, a mechanism that is solved, a transaction that does or does not bankrupt the agent. Video generation has no verifier, and scalar drift is what the reward does in its absence. A rubric generated per query is a way of narrowing the reward model's question without narrowing its domain, which is the same trade Learning to Discover Interesting Mathematics makes by reducing "interesting" to a ratio and Coding Agents for Generalized Task and Motion Planning Problems gets free from a simulator. Three papers captured the same day, three ways of manufacturing a verifier where none exists.
The claim is asserted, not quantified, and that is a gap in the capture. The
snapshot carries no table and no percentages — "state-of-the-art" and "mitigates scalar drift" are the whole of the results text, so no figure from this paper is
quotable as a score, including the drift reduction the architecture exists to
deliver. Authors and affiliation are not stated in the snapshot and arxiv.org
is blocked from this run.
The structural objection: the rubric generator is itself trained, so the drift has been moved rather than removed — a generator can drift across prompts exactly as a scorer can. The answer offered is the joint optimisation against human ratings, which makes human ratings the anchor of last resort, and their number, source and scale are not stated. No downstream result is reported either: a better reward model is not yet better generated video.
State of the Art (2026-09-27)
This page has been accumulating fixes to the reward signal. Today's paper removes the step where a reward gets written by hand at all.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms (2026-09-23) identifies the defect as an ordering one, and the whole contribution is reversing it. Existing pipelines "construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc." VHD-Play instead samples and solves a mathematical model first, then has a corpus-grounded setter render that solved model's decision process as stateful tools — so the executable dynamics and the trajectory-scoring reference are inherited from the same solved model rather than written separately and reconciled (source).
| Measurement | Figure |
|---|---|
| Environments generated | 3,300, at a few cents each |
| Qwen3.6-35B-A3B, mean agentic score, five-family diagnostic | 0.204 → 0.815 |
| Generalisation | held-out instances from all three training families, plus eight unseen mechanism families |
| External transfer | general function calling, travel planning, 365-day e-commerce |
| E-Commerce Bench | completes every run without bankruptcy, exceeds Qwen3.7-Max |
| The ablation is the result, not the score. Comparing written-out problems against | |
| stateful versions that reveal or hide their parameters, the paper reports that **most | |
| of the learnable gap lies in stateful interaction rather than in the underlying problem | |
| solving.** That is a claim about what agentic RL is teaching: not the mathematics, but | |
| operating a stateful interface — and it is the most direct evidence this page holds | |
| for why agentic benchmarks and reasoning benchmarks come apart. |
0.204 → 0.815 is a 4× move on an undefined scale. Nothing read states what the agentic score measures, its range, or whether 0.815 is near a ceiling, so it is not compared with anything else here.
Its pairing with PACT: From Credit Assignment to Critic Alignment is close to exact and the two do not cite each other. PACT proves what a correct credit signal is, given an environment — three conditions uniquely determining token-level credit. VHD-Play makes the environment so the credit signal is inherited. One is a theory of the reward; the other deletes the step at which a human writes one.
And it pairs from the opposite side with ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds, captured the same day: the same scarce resource — a verifiable unfamiliar mechanism — manufactured once to train on and built once to test with, by two groups who did not read each other. The irony is worth recording: VHD-Play's environments come from solved mathematical models, which may sit closer to pre-training data than "hidden dynamics" implies, and nothing read tests for it — the exact control ExplorationBench was built around.
A second entry, on ordering rather than on rewards. Rufus-Air: An Open LLM Post-Training Recipe documents eight serial post-training stages and draws from them the rule that reward reliability is what should order the stages — hard verifiable rewards first, softer judge-based signals last, because a later stage inherits the policy an earlier one left. It is the plainest statement of stage ordering this page holds, and it comes with no benchmark figure of any kind (source).
State of the Art (2026-09-26)
-
Token-level credit gets a uniqueness proof, and the algorithm it motivates gains least where the proof matters most (2026-09-22): PACT: From Credit Assignment to Critic Alignment opens by stating that token-level credit "lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear" — which is this page's condition, written down. It formulates three regularity conditions, Completeness, Prefix Consistency and Neutrality, and proves they uniquely determine token-level credit. Three consequences are then read off that definition: an ideal teacher in On-Policy Distillation acts as an implicit critic, yielding an expected policy gradient proportional to the one token-level credit induces; response-level RLOO signals match the expected policy-gradient contribution despite coarser granularity; and in GAE, intermediate critic errors can become comparable to the underlying credit itself. The paper further establishes approximate credit sparsity under bounded outcome rewards. PACT adopts an Actor-then-Critic update order, which is what makes importance-sampling correction on critic training available and aligns the critic with the updated policy. Reported: 72.87% average accuracy on four agentic mathematical-reasoning benchmarks — +8.80 over GRPO, +13.16 over PPO — and 67.4% pass rate on SWE-bench Verified, +2.4 over PPO, +2.0 over GRPO, +3.8 over SAO (source). Why it belongs at the top: every entry below defends a credit-redistribution scheme with a benchmark, and this is the first object against which such a scheme can be wrong rather than merely behind. The OPD-as-implicit-critic result lands directly on the on-policy-distillation cluster this page has been accumulating since 2026-09-02 — that cluster's open question is what the teacher's signal is actually buying, and this answers it in the form of an equivalence rather than an ablation. The GAE result has the most immediate teeth: a widely used estimator can be dominated by its own noise in exactly the long-horizon regime agentic RL runs in. The results argue against the motivation and the paper does not reconcile them: the margin over GRPO is 8.80 points on agentic maths and 2.0 on SWE-bench Verified — four times smaller on the benchmark where credit is spread over the longest horizon, which is the opposite of what a theory of long-horizon credit assignment predicts. Held with a stated gap: no base model, parameter count or training compute; the four maths benchmarks are unnamed; SAO is named only as a baseline and not expanded; no seed count or variance figure. And the 67.4% carries an asterisk the paper does not acknowledge — Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?, captured 2026-09-25, consistently degrades SWE-bench Verified performance by instantiating the repository at evaluation time
-
Environment synthesis drops its last human artefact: not issues, not commits, just the source (2026-09-18): CodeMidas: Scaling Agentic Coding RL Environments from Code Itself turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input — explicitly against the prevailing practice of mining issues and commits, which bounds the extractable task set to what somebody wrote down. Agentic compute is spent at every stage of construction: agents explore the code to formulate behavioral specifications, construct tests grounded in execution of the original code, then validate and filter through execution checks and repeated solution rollouts. Yield: 5,545 training tasks from 3,185 open-source codebases, 23 languages, 15 technical domains. Training MiMo-V2.5 with GRPO on them gives DeepSWE +11.7%, ProgramBench +17%, Terminal-Bench v2.1 +8.5%, with an ablation reporting that more high-quality tasks keeps helping and trajectory analysis reporting more codebase exploration and more diverse self-verification. This is SPADE's "environment as a learnable object" with the generator grounded in code that already runs — SPADE writes environments and their verifiers from scratch, EnvHarness reshapes a trusted one, CodeMidas derives both from an artefact whose behaviour is executable and therefore checkable without a model's say-so. What it does not close: "all five benchmarks improve" is stated while only three are named with figures; every number is a relative gain with no absolute score, so none is comparable to the Terminal-Bench 2.1 figures this wiki holds elsewhere; only one base model was trained; and no contamination check is reported between training codebases and benchmarks built from the same population. The filter is its own bias — "repeated solution rollouts" keeps tasks the current model can sometimes already solve (source)
-
The tenth OPD result removes the third ingredient in eight days, and the three removals together relocate the bottleneck (2026-09-05): Rethinking On-Policy Distillation of Large Language Models II: One Training Example trains on a single query and reports that one-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. The mechanism offered is state coverage — the fraction of the states full-data OPD visits that a query set's rollouts also reach: one query already reaches 71.5%, most of it within the first 100 steps, and 16 semantically distinct queries reach 98.9% and match full-data training. The result extends to multi-teacher OPD (16 diverse queries per domain match full-data MOPD), and survives a stress test where content-light templates and off-domain WildChat queries approach the real-query baseline — so task content and induced state coverage come apart. Meanwhile alignment slows at a similar pace whether training on one query or the whole dataset, and even a fixed state set takes hundreds of steps to absorb. The authors' summary: OPD is "data-overfed but algorithm-starved" (source). Why this belongs above the ninth: the three most recent results each delete a different thing this page assumed was carrying OPD, and each reports little loss — Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement deletes the teacher's content (a fixed negative advantage matches teacher-provided ones),
2608.29846IDA-OPD deletes the full-vocabulary signal (sampled-token log-probability suffices), and this deletes the data. The 09-04 entry recorded the first two as an unresolved tension — one paper optimising a signal the other says can be discarded. This does not resolve that tension; it shrinks it. If the student's absorption rate is independent of how much supervision is available, both arguments are about a part that is not rate-limiting. The pairing is this wiki's reading, labelled as such: none of the three cites the others. Held with a stated gap: "most of full-data OPD's gain" carries no number, the domains and model families are unnamed, and every figure is on query-answer training rather than at agentic horizons. The title says "II" and nothing read identifies paper I -
A rubric result worth recording without a page, because it works where this page's whole OPD lane cannot (2026-09-05):
2609.04094DRACO operates in the outcome-blind setting — no programmatic checker, which is most long-horizon agent domains. Multi-criteria rubrics are the usual substitute but are scored once per trajectory, and one scalar is a poor signal across tens of steps. DRACO generates rubrics dynamically during training to track the policy's evolving capability, scores them once per completed trajectory, then redistributes that judgment over the steps responsible for the annotated rubrics to produce per-step advantages in GRPO — closed-form, with no trained attribution module. Reported: +15.9 points over the base model on AppWorld and +5.3 over GRPO trained with a sparse ground-truth reward, without using any verifier itself; +5.3 on out-of-domain Tau-Bench even without a frontier judge, beating both ground-truth-reward training and other rubric-based settings. Code atgithub.com/IBM/draco(source). Recorded here rather than given a page, following the precedent set for Cliff and IDA-OPD: it is a credit-assignment variant on a mechanism this page already tracks. What makes it worth the lines: beating a sparse ground-truth reward while using no verifier is the counter-intuitive row, and it points at the same conclusion as the OPD cluster above — the signal's shape across steps is doing more work than its source -
The ninth OPD result answers the eighth's implicit question — if the teacher's content is not doing the work, what is the teacher's signal actually costing you? (2026-09-04):
2608.29846Influence-Directed Distillation names the failure as diversity distillation failure — the student's pass@1 improves while pass@k plateaus, so it does not inherit the teacher's diversity. Its diagnosis is a signed first-order proxy, First-Order Local Entropy Influence, which decomposes each update's entropy effect into the teacher–student log-probability gap and the student's local probability structure, and links entropy contraction to negative-influence positions. IDA-OPD then keeps entropy-expanding updates and replaces entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability rather than full-vocabulary Forward-KL. Reported: pass@k consistently improved, matching the strongest teacher-informed methods at strictly lower cost, with vanilla OPD's pass@1 broadly maintained (source). Why it sits directly beneath the eighth: Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement reports the teacher's content is not what makes OPD work and replaces it with a constant; this reports the teacher's signal shape is actively collapsing entropy and fixes the shape instead. They are not the same claim and they are not compatible in the obvious way — if a fixed negative advantage matches teacher-provided ones, IDA-OPD's careful use of the teacher's sampled-token log-probability is optimising something the other paper says can be discarded. Neither cites the other; both landed inside three days. Held at abstract confidence: IDA-OPD's abstract carries no absolute figures, no benchmark name and no model name, so every result above is a direction without a magnitude. A third paper from the same snapshot attacks the same seam from the reward side:2609.02817Cliff uses an off-the-shelf LLM to find the first mistake in a rollout and converts it into token-level advantages — positive prefix, negative suffix — reporting +15% over on-policy distillation and +7% over standard GRPO across 12 scenarios, "even with teachers of modest capability". Recorded here rather than given a page: it is the third variation on one week's theme and the cluster already carries eight -
The eighth on-policy-distillation result in this cluster is the first that says the teacher is not the reason it works (2026-09-02): Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement measures teacher supervision during OPD training and reports it is substantially noisy, with noise prevalence rising with teacher scale — and that the student is insensitive to that noise, converging to comparable performance whether it is kept or removed. Two further probes: learning concentrates on low log-probability tokens, and a single fixed negative advantage matches teacher-provided ones. The mechanism the authors are left with is suppression of low log-probability tokens, which requires no teacher. Their supervision-free replacement, OPSA (entropy-adaptive negative advantages), reports +35.41 Avg@32 on AIME24 over base Qwen3-1.7B (a 263% relative gain), Pass@32 more than doubled across all three benchmarks, and +16.77 Avg@32 over OPD itself. Why it belongs at the top of this page: seven results above spend OPD's dense token-level signal on something — security, RLVR gating, multi-teacher budgets, test-time labels — and every one of them assumes the signal carries the teacher's content. This is the first to test that assumption directly and report it does not. It is also the sharpest possible form of the page's data-efficiency problem: if a fixed constant matches a teacher, the teacher's inference cost bought nothing. It cites none of the seven, and none of them cites it. The scope is narrow and the abstract admits most of it: the base is Qwen3-1.7B, the "all three benchmarks" are not named (only AIME24 is), and nothing read separates "suppressing tail tokens" as a capability gain from the same description applied to a sharpened output distribution (source)
-
Three papers in four days say GRPO's limits are a design choice, and they disagree about which one (2026-08-29): Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351) makes the third and the most direct — it argues Evolution Strategies is not the memory-efficient-but-weaker alternative to GRPO it is usually filed as, but a distinct post-training paradigm. The claim: GRPO exhibits entropy collapse and ES does not, so ES "improves Pass@1 while attaining higher Pass@K", with a theoretical account resting on verifier-projected Jensen-Shannon diversity across the ES population; a sequential GRPO-ES schedule is proposed to take Pass@1 from one and Pass@K from the other. Set against the two already on this page, the disagreement is the interesting part: BPCO (
2608.23566) says the fixable choice is GRPO's critic-free premise, Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311) says it is where the KL regulariser sits, and this says it is the optimizer itself. Why it belongs here and not only on the paper page: every other result on this page competes on the rollout budget, and this one competes on coverage — Pass@K is a measure of how much of the pretrained model's reasoning remains reachable after post-training, which is the quantity Post-Training Scaling implicitly assumes is preserved. The second finding travels further than the comparison: functional sparsity, that ES's gains come from a sparse subset of large-magnitude updates "despite substantial whole-model parameter drift", with held-out evaluations showing no necessary catastrophic forgetting — if that holds, measuring how far weights moved says nothing about how much function changed, and several claims in this wiki reason from exactly that. The abstract carries no absolute figures, names no benchmark and names no model, so every comparison above is a direction without a magnitude (source) -
The label comes off entirely, and the multi-teacher seesaw gets diagnosed (2026-08-28): three more on-policy distillation results, taking the count to seven in four snapshots. TTPO: Test-Time Policy Optimization (arXiv:2608.27448) removes the ground-truth label, which is what has kept every method on this page out of test-time training. Majority-vote pseudo-labels are fragile — "an incorrect vote corrupts the teacher and misleads every token" — so TTPO exploits an asymmetry: rollouts that disagree with the pseudo-label "are typically wrong regardless of whether the vote itself is correct". Agreeing rollouts are distilled via OPSD, disagreeing ones penalised with Grouped RL, with token-level selection down-weighting converged positions and penalising only confident errors. Without any labels it matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, and yields +25.2% to +36.4% without thinking. The regime it does not test is the one that matters: where the majority is systematically wrong, the minority rollouts are the correct ones and TTPO penalises exactly them. The asymmetry is an observed property of the studied distribution, not a theorem.
Two papers in the same snapshot attack the multi-teacher seesaw this page records below as an open problem. Open-MOPD (
2608.19098) builds a controlled M-OPD benchmark on SmolLM3-3B-Base with oracle routing, isolating capability integration from routing ambiguity, and finds standard M-OPD captures only 35.6% of the available headroom relative to a domain-routed oracle ensemble. Its diagnosis contradicts the intuitive one: the failure "stems not from gradient conflict, but from a severe misallocation of the token-level optimization budget", driven by structural sequence-length disparities across domains, convergence drift from non-uniform learning rates, and multi-step reward staleness from asynchronous updates. Token-share balancing, gap-aware dynamic budget allocation and student reward refresh lift headroom recovery 35.6% → 83.4% in a single deployable student, with the end-to-end recipe, trajectories and evaluation suites open-sourced "on an academically accessible hardware budget". D³-MOPD (2608.24987) attacks the same waste from the scheduling side: a zero-overhead off-process watcher reuses the per-domain reverse-KL signal already produced during training to adapt the domain mixture online, closing 97% of the average student-to-teacher gap against 63% for vanilla MOPD on a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, reaching peak performance with "approximately 3times reduction in rollout steps" and surpassing the specialist teachers on three of seven benchmarks. Neither is given a page, following the precedent set for OPDVR and BPCO: both are recipe-level results on an existing mechanism this page already tracks. Neither cites Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647), whose origin-relationship finding is the strongest existing explanation for the seesaw they are fixing empirically (source) -
On-policy distillation appears four times in two snapshots, and the common move is token-level supervision replacing trajectory-level reward (2026-08-27): the 08-25 controlled study above established what OPD transfers; this snapshot shows the field spending it. SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation (arXiv:2608.21500) points it at security — a rollout on an injected input is scored token by token by the initialization model given the clean input, so the model learns which tokens the injection produced. Reported ASR against the adaptive attacker PISmith: 94.0% → 9.0% against Meta-SecAlign's prior SoTA, on Qwen3.6-27B, weights released. OPDVR (
2608.24696) combines OPD with RLVR without adding a hyperparameter: it reformulates sampled-token OPD's implicit reward by trajectory correctness and applies ReLU gating so correct trajectories take non-negative rewards and incorrect ones non-positive — which turns sampled-token OPD into a proper RLVR method combinable with GRPO, reported to beat standard OPD on six reasoning benchmarks. A third (2608.24646) carries on-policy self-distillation into diffusion models. The unifying diagnosis is the one SecOPD states outright — sequence-level feedback "prevents the model from learning precisely which output tokens are insecure" — and the same argument underwrites OPDVR's dense-signal half. Neither OPDVR nor SecOPD cites the 08-25 generalization study, whose finding constrains both: OPD's reach is set by the origin relationship, and SecOPD's teacher is the student's own initialization, the maximally same-origin case (source) -
A critic-based recipe is reported to match group-relative methods at one sample per prompt (2026-08-27): BPCO (
2608.23566) revisits the reason GRPO exists — critics are unstable — and reports a recipe that stabilises one: DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive GAE. Because the critic runs only during training it can be conditioned on reward-defining information hidden from the policy — a reference answer or grading rubric. Across mathematical reasoning from 1.5B to 30B-A3B, it is reported to match or exceed a group-based baseline while sampling one response per prompt. Why it belongs here: every agentic RL result on this page pays the group-sampling tax, and the rollout budget is the binding cost. No agentic benchmark is reported — the evaluation is mathematical reasoning — so whether it survives long-horizon credit assignment is untested (source) -
Environment generation gets its scale number, and the authoring turns out to be learnable too (2026-08-26): AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale (arXiv:2608.20634) inverts the construction order this page has recorded twice — SPADE writes an environment for a task, EnvHarness reshapes one; AgentMercury instantiates a persistent world (entities, services, tools, state, executable cross-service invariants) from a business scenario and lets tasks emerge from it. 4,783 executable environments, 14 industries, 50 countries, used as RL substrate. The transfer claim is the one to check: Qwen3.5-4B 12.3 → 15.7 on EnterpriseOps-GYM and 45.9 → 56.0 on AIME26 — +10.1 points on competition mathematics from training on business workflows, with the environments stated to be generated "without targeting the evaluation benchmarks". The second experiment is the more consequential: fine-tuning Qwen3.5-35B-A3B on construction traces raises executable-world authoring success on held-out scenarios from 3.3% to 83.3%. If authoring verified environments is itself learnable, this page's standing bottleneck moves from human-authored environments to compute. No contamination analysis is mentioned, "generated without targeting" is a claim about intent rather than about distribution overlap, and 83.3% success is never defined — does the world merely run, or does it hold its invariants? (source)
-
The stability–exploration trade-off is argued to be an artefact of where the regulariser sits (2026-08-26): Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311), from an Alibaba / Qwen AI Lab org repository, keeps the drift control and gives back the exploration budget by moving the constraint from the action side to the input side. Query-KL bounds how far the distribution over training queries induced by the current policy drifts from its pre-RL reference; because the QKL gradient flows strictly through the query likelihood and the response score function does not appear in the term, it "exerts no direct gradient pressure on the response distribution". Plus a dataset-static, reference-derived per-query weight. It replaces Policy-KL in GRPO/PPO/REINFORCE pipelines with no additional forward passes, reported to give stronger accuracy and "substantially more stable behaviour under high-temperature decoding and long-horizon training" on six mathematical reasoning benchmarks. Why it belongs here rather than in a general RL note: high temperature and long horizons are the regime agentic RL runs in, and the usual remedy is to leave it. The abstract carries no numbers, names no benchmark and names no baseline coefficient, and nothing read demonstrates that unbounded query drift causes the instability it is blamed for — only that ERPO bounds it (source)
-
On-policy distillation transfers behaviour, and how far it reaches is set by a relationship rather than by teacher quality (2026-08-25): Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647) is a controlled study varying one generalization factor at a time, and it reports three things this page's OPD lane assumed rather than measured. What transfers is the teacher's reasoning behaviour, not its answers: training difficulty barely matters and problems the teacher never solves are still useful — which is the mechanism behind Direct-OPD's weak-model rollouts, now stated as a finding. Same-origin teacher/student pairs carry across languages, reasoning horizons and other domains; cross-origin pairs mostly fit the trained distribution. And the reach is not steerable: because routing prompts to domain experts cannot confine a teacher's influence, multi-teacher OPD produces a mixture-dependent seesaw among the teachers' capabilities rather than their union. The abstract carries no numbers at all, and "origin" is never defined — same corpus, same family, same lab are three different claims (source)
-
The training environment itself becomes a learned component — now by augmentation, not only generation (2026-08-22): EnvHarness: Awakening Static Worlds for Agent Learning (arXiv:2608.19880) keeps the base policy's environment but wraps it in a programmable plug-in layer that reshapes behaviour to target the policy's diagnosed weaknesses (synthesised by an automated component, EnvRigger, from black-box trajectory observation) while retaining the original verifier — reported to give a superior RL optimization signal and enable continuous co-evolution of policy and environment, up to +9.0 points on held-out instances. This is SPADE's "environment as a learnable object" with the trust problem solved from the other side: SPADE writes environments (and their verifiers) from scratch; EnvHarness reshapes a trusted one and never replaces its verifier — directly addressing the open problem below that SPADE's self-authored verifiers are an undefended reward-hacking surface (source)
-
Harnessed agentic RL — the deploy-time harness now participates in training (2026-08-18): Agent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528) names the regime and ships a ~3,500-line framework for it. Arbitrary agents connect to training through an LLM endpoint proxy, so the harness owns the environment interaction loop and the trainer observes only LLM request/response pairs. The paper states four other frameworks — verl Uni-Agent, AReaL 2.0, slime, Polar — have adopted the architecture. Reported: Qwen3.5-9B 41.8% → 56.4% on SWE-bench Verified from 6K examples. Consequence: a model post-trained this way is fitted to its harness, which is the mechanism behind ClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798)'s finding that such a model stops being harness-agnostic. The harness it used for the 56.4% is not named (source)
-
The competing bet: delete the credit-assignment machinery rather than fix it (2026-08-18): Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310) argues evolution strategies suit long-horizon agents better than RL — full-parameter optimization at inference-level GPU memory, and trajectory-level parameter attribution with no reward decomposition across the horizon. Reported: +6.69% over a No-Skill baseline on WebArena-Lite for Qwen-3.5-27B, and its matched baseline beaten in 28 of 36 test-time heuristic-design settings. No head-to-head against agentic RL is reported, and no rollout count is published, so the efficiency claim is memory, not compute (source)
-
The harnesses get named, and so does the integrity check nobody else reports (2026-08-18, recorded 08-21): LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393) trains Qwen3.5-35B-A3B with GSPO inside three unmodified production harnesses — OpenHands SDK 64.0% → 70.4%, Claude Code 62.4% → 68.2%, OpenCode 57.2% → 66.6% on SWE-bench Verified. It supplies exactly what Agent Lightning withheld: the harness identity, three times over. Its mechanism is in-process LLM proxying plus trainer-side log-probability recomputation, which survives the harness compacting or re-serialising context mid-rollout — and it reports the quantity that decoupling destroys, rollout–training probability correlation above 0.99. No benchmark delta reveals a broken correlation, so a harness-training paper that omits this number has not shown its optimization was valid (source)
-
Removing the label, from two directions, on one day (2026-08-21): Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (arXiv:2608.17253) optimizes multiple decoupled models sharing no parameters with rewards derived from peers, and identifies the failure of self-rewarding RL as an error-correlation problem rather than a bad-judge problem — increasing cohort diversity (families, sizes, rephrased samples) reduces the correlated errors that drive collapse. Reported 3.0–8.6% across seven text benchmarks and 2.3–7.2% across four multimodal ones, with no ground-truth labels, stated to match or surpass supervised methods. SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197) removes the same dependency from the task side: one LLM writes complete executable training environments with
reset/step, targeted by a regret estimate — the gap between the learner's reward with and without privileged hints — giving +5.3 average over the strongest fixed-environment baseline at 30B, +5.7 on BFCL-v4 multi-turn and +13.9 on ACEBench-Agent. Co-RL generates the reward without labels; SPADE generates the tasks. Neither cites the other (source) -
Anthropic AAR (2026-04-14): 9 instances of Claude Opus 4.6 worked on an alignment research problem (weak-to-strong supervision) for one week in an agentic RL setup → PGR 97% vs. 23% for human researchers. The first empirical case of RL agents replacing humans on a real-world research problem. → Automated Weak-to-Strong Researcher (AAR)
-
Self-Distilled Agentic RL (2026-05-16 HF Daily #3): Self-Distilled Agentic RL
-
Direct On-Policy Distillation / Direct-OPD (2026-07-06, Tsinghua AIR + ByteDance Seed): Run RL on a cheap weak model, then transfer only the RL-induced policy shift (
Δ = log π_T − log π_{T,ref}) to a large model using on-policy distillation. Cuts the cost of deploying RLVR gains at frontier scale. A practical cost lever for the Anthropic AAR-style workload where frontier RL is expensive. → Weak-to-Strong Generalization via Direct On-Policy Distillation -
Related prior work: ReAct, Reflexion, AutoGPT family (page TBD)
Open Problems
- Multi-step credit assignment — Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310) proposes sidestepping it entirely via trajectory-level attribution, rather than solving it
- Reward hacking in tool-use environments — LEGO-RL spends one of three pillars on stage-wise defenses against it, and SPADE has the sharper version of the problem: its environments write their own verification code, authored by a model sharing weights with the agent being trained. EnvHarness: Awakening Static Worlds for Agent Learning (arXiv:2608.19880) is the counter-design: retain the original, trusted verifier and only reshape the environment around it — so the reward signal is never authored by the party being optimized
- Data efficiency (online RL is expensive) — Direct-OPD partially addresses via weak-model rollouts. And a training-free branch now competes for the same gains: FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills (arXiv:2607.21596) and Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses (arXiv:2608.08466) both leave the model frozen and evolve the scaffold instead, reporting double-digit gains with no gradient step at all — see Agents (LLM Agents). Nothing read compared the two routes on the same task, which is the experiment that would say whether agentic RL is buying something a rewritten harness cannot. Run on 2026-08-31 by WHALE: A Simple Recipe for Joint Harness-Weight Optimization, and the answer is "it depends on the domain" — weight-only, harness-only and alternating, on the same tasks with Qwen3.5-2B/4B: harness search matches peak weight-only accuracy with far fewer rollouts on SearchQA, and improves math accuracy only after a weight update. Alternating beats all three baselines by 4.15–24.38 pp. So neither route dominates, the bottleneck moves between domains, and a fixed budget split is wrong somewhere. The question is now bounded rather than closed: 2B/4B is not frontier scale, all three domains are verifiable, and nothing read tests the alternation where the reward is outcome-blind
- Multi-teacher distillation is not additive — Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647) reports a mixture-dependent seesaw, because routing cannot confine a teacher's influence. Two 2026-08-28 results attack this empirically and reach 83.4% and 97% headroom recovery, but by fixing budget allocation and scheduling rather than by addressing the routing argument — Open-MOPD explicitly uses oracle routing to remove routing ambiguity from its benchmark, which is the variable the original finding turns on. The problem is therefore better bounded than before and not closed
- Long-horizon stability
Key Papers
-
Automated Weak-to-Strong Researcher (AAR) — Anthropic AAR, 2026-04-14 (97% PGR, weak-to-strong supervision)
-
Weak-to-Strong Generalization via Direct On-Policy Distillation — Direct-OPD, Tsinghua AIR/ByteDance, 2026-07-06 (cost-efficient RL transfer)
-
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning — SEED, 2026-07-14 (on-policy distillation for agents)
-
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning — Ring-Zero, Ant Group, 2026-07-12 (RLVR at a trillion parameters; five emergent behaviours)
-
Agent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528) — Agent Lightning v1.0, 2026-08-18 (harnessed agentic RL; SWE-bench Verified 41.8% → 56.4% on Qwen3.5-9B)
-
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310) — Agentic ESOpt, 2026-08-18 (evolution strategies instead of RL for long-horizon agents)
-
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393) — LEGO-RL, 2026-08-18 (harness-native RL in three named production harnesses; 0.99 rollout–training correlation)
-
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (arXiv:2608.17253) — Co-RL, 2026-08-19 (peer reward across a diverse cohort; unsupervised reasoning without labels)
-
SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197) — SPADE, 2026-08-19 (the training environment itself as a learned component)
-
EnvHarness: Awakening Static Worlds for Agent Learning (arXiv:2608.19880) — EnvHarness, 2026-08-20 (reshape a static environment while keeping its verifier; policy–environment co-evolution)
-
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647) — 2026-08-25 (controlled study: OPD transfers behaviour not answers; same-origin transfers broadly, multi-teacher seesaws)
-
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale (arXiv:2608.20634) — AgentMercury, 2026-08-26 (4,783 scenario-grounded environments; environment authoring 3.3% → 83.3%)
-
Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311) — ERPO, Alibaba, 2026-08-26 (Query-KL on the input side replaces Policy-KL)
-
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation (arXiv:2608.21500) — SecOPD, 2026-08-27 (token-level OPD as a prompt-injection defence; PISmith ASR 94.0% → 9.0%)
-
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351) — 2026-08-27 (ES avoids GRPO's entropy collapse; higher Pass@K, and gains carried by a sparse subset of updates)
-
J-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero Data — J-Zero, 2026-08-27 (Challenger–Solver–Judge co-evolution from zero data; the Judge's preference labels come from how a response was produced, not from its own scores; +4.2 verifiable / +8.0 unverifiable, still improving at ten iterations where baselines degrade after two — the first result here claiming movement on the unverifiable side of the boundary AI Alignment drew in August)
-
Rethinking On-Policy Distillation of Large Language Models II: One Training Example — 2026-09-03 (one query recovers most of full-data OPD; state coverage 71.5% at one query, 98.9% at sixteen; "data-overfed but algorithm-starved")
-
WHALE: A Simple Recipe for Joint Harness-Weight Optimization — WHALE, 2026-08-31 (weights and harness optimised by alternation; +4.15–24.38 pp, and either component can be the bottleneck depending on domain)
-
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments — Terminal-Universe, 2026-09-03 (agent trajectories replayed back into executable environments; the supply side of environment-scaled post-training)
-
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL — ContextPilot, Tencent, 2026-08-28 (action-level credit for context edits via branch sampling at high-entropy decisions, replacing trajectory-level reward; adds planning, long-term memory and soft offloading tools; no figure published in the abstract)
-
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms — Verifiable Hidden Dynamics Play, 2026-09-23. Environments generated from solved mechanisms so dynamics and scoring are inherited rather than aligned post hoc; 3,300 environments, 0.204 → 0.815, and the finding that most of the learnable gap is stateful interaction rather than problem solving (source)
-
Rufus-Air: An Open LLM Post-Training Recipe — Rufus-Air, 2026-09-24. Eight-stage recipe whose stated ordering principle is reward reliability; no figure published (source)
Sources
Referenced by
Sources
- sources/arxiv/2026-10-04/2609.35259-distillation-dynamics.md
- sources/papers-daily/hf-daily-2026-09-30.md
- sources/papers-daily/hf-daily-2026-09-28.md
- sources/papers-daily/hf-daily-2026-09-22.md
- sources/papers-daily/hf-daily-2026-09-05.md
- sources/papers-daily/hf-daily-2026-09-04.md
- sources/papers-daily/hf-daily-2026-09-02.md
- sources/papers-daily/hf-daily-2026-09-01.md
- sources/papers-daily/hf-daily-2026-08-29.md
- sources/papers-daily/hf-daily-2026-08-28.md
- sources/papers-daily/hf-daily-2026-08-27.md
- sources/papers-daily/hf-daily-2026-08-26.md
- sources/papers-daily/hf-daily-2026-08-25.md
- sources/papers-daily/hf-daily-2026-08-21.md
- sources/papers-daily/hf-daily-2026-08-20.md
- sources/arxiv/2026-05-16/2605.15155-self-distilled-agentic-rl.md
- sources/blogs/anthropic-2026-04-14-automated-alignment-researcher.md
- sources/arxiv/2026-07-06/2607.05394-weak-to-strong-opd.md
- sources/papers-daily/hf-daily-2026-09-27.md