AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/post-training-scaling.md

Post-Training Scaling

conceptupdated 2026-10-01created 2026-08-21

Definition

The claim that the next increment of model capability is bought after pre-training is finished — by scaling RL, agentic training environments and the harness a model is trained inside — rather than by adding parameters or pre-training tokens.

Stated as a scaling law by a lab CEO for the first time on 2026-08-20, and stated as a distinction between capability types:

CapabilityWhat it is reported to prefer
Memorizationmore parameters
Reasoningmore post-training data and effective depth
Advanced skills (named example: finding software vulnerabilities)not total parameter count, once a knowledge-holding threshold is reached
Attributed to Jie Tang, CEO of Z.ai, reported by Latent Space's
AINews digest under the title "Death of Params"
(source).

This page collects a claim, not a settled result. The vendor framing and the research results below are recorded separately and deliberately: one is a lab's account of its own model, the others are independent papers that happen to point the same way.

Why It Matters

Because it is the first time the "no new pre-training" pattern has been named by someone shipping models rather than observed by this wiki. The Weekly Synthesis — W33 (2026-08-10 → 2026-08-16) synthesis recorded that four releases in one week trained nothing new — GLM-5.3, Gemini 3.7 Flash, DeepSeek V4-Pro-0813 and Qwen's post-training column. That was a pattern read off release notes. Jie Tang's claim is the same pattern asserted as a scaling law with a mechanism, and it comes with a model offered as the experiment.

If it holds, three things this wiki tracks change meaning:

  1. A parameter count stops being a capability signal. Every model page here carries one, and the comparison pages are built from ## Spec tables in which it is the most legible row.
  2. "Open weights" stops transferring capability. GLM-5.3's base is GLM-5.2's, and GLM-5.2's weights have been MIT-licensed since 2026-06-16. Anyone could already have the base; what they cannot have is the post-training. Open-Weights Policy Fight has treated open weights as the axis of the open/closed dispute — this makes the released artefact the less valuable half.
  3. The harness enters the training loop, so a benchmark number stops being a claim about a (model, harness) pair at evaluation time and becomes one baked into the weights. That consequence is developed on Eval Harness Configuration.

The counter-consideration, kept on the page rather than resolved: "post-training scaling" describes where the effort went, not that pre-training has stopped paying. Gemini 4's pre-training run was confirmed on 2026-07-21 as Google's "most ambitious yet", and nothing read reports its results. A field harvesting post-training gains while a frontier pre-training run is still in flight is not evidence that the run will fail.

State of the Art (as of 2026-10-01)

Power laws for on-policy distillation, and the result cuts against this page's central claim. Scaling Properties of Same-Family On-Policy Distillation — the top entry in today's HuggingFace snapshot at 194 upvotes — fits scaling laws to same-family OPD across weak-to-strong, same-base and strong-to-weak pairs (source).

ClaimFinding
Transfer regimeheld-out accuracy G rises approximately linearly in sqrt(KL(π_θ ‖ π_ref)) from the student's initialization
Weak-to-strongstudent's peak exceeds its teacher's own score in every observed pair
Teacher scaleG_peak improves with teacher scale only up to roughly the student's scale
Teacher scoreat matched gold score, smaller teachers transfer better
**This page's argument is that capability increasingly comes from the
post-training stage rather than from parameters** — Z.ai's "death of params"
framing, and the 53 → 60 index move on an unchanged base. This paper says the
product of that stage is cheaply movable, and most cheaply by the smallest
thing that has it. If post-training is where the value is, and post-training's
output transfers weak-to-strong past the teacher's own ceiling, then the moat is
thinner than the thesis implies.

The sqrt reverse-KL parameterisation is the reusable part. It gives OPD an x-axis measurable from the student alone, with no reference to the teacher — which is what makes a weak-to-strong comparison meaningful rather than circular.

No absolute figure, no model family, no parameter count appears in the abstract this snapshot carries. These are the shapes of the laws, not points on them, and arxiv.org is blocked from this sandbox so the fitted constants cannot be read.

It lands the day after the second lab-level distillation accusation, and the adjacency is recorded on Adversarial Distillation — the paper is about a lab distilling within its own family and says nothing about unauthorised use. Mechanism, not evidence.

State of the Art (as of 2026-09-27)

The first fully documented post-training recipe this page holds, and it publishes no numbers.

Rufus-Air: An Open LLM Post-Training Recipe (2026-09-24) documents eight serial stages on GLM-4.5-Air-Base (106B-A12B) — SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, RLHF — and states it documents the data, reward design, infrastructure, stage order and stagewise results needed to reproduce the recipe (source).

The constraint is the contribution: built on open-source components and public data, much of it used as released, with no new human annotation and no in-house distillation teacher. Almost everything on this page is a lab describing what post-training bought its own frontier model with the recipe withheld. This is the opposite document.

Four stated findings, of which the third is the one with teeth:

  1. Diverse, high-quality SFT establishes a strong capability floor.
  2. Difficulty filtering keeps RL prompts within a productive learning range.
  3. Reward reliability provides a practical principle for ordering stages.
  4. Infrastructure and engineering choices are part of the recipe, not just an implementation detail.

Finding 3 restated: put the stages whose rewards you can trust first, because a later stage inherits the policy the earlier one left. The eight-stage order is consistent with it — verifiable reasoning and coding RL before judge-scored RLHF.

And then there are no figures. "Improves over the official GLM-4.5-Air post-trained release" and "competitive with similarly sized open models" name no benchmark, no margin and no comparator. This page records that as an absence rather than paraphrasing it as a result: a reproducibility paper whose headline claims cannot be checked is a recipe, not a finding.

Context that bears on how much the comparison is worth: this repo's captured Artificial Analysis table has no GLM-4.5-Air row. Its Z AI entries are GLM-5.3 (max) 45, GLM-5.3-Flash 42, GLM-5.3 (low) 34 (source) — the base being improved on is two minor generations behind what that publisher currently lists.

Why the "no distillation teacher" line matters here. Anthropic's September threat-intelligence report accuses several Chinese labs of distilling Claude, and Alibaba / Qwen AI Lab carries two unreconciled figure sets for a campaign said to have trained Qwen 3.5–3.7. A recipe reaching "competitive with similarly sized open models" without a frontier teacher would be a data point on what the teacher was worth. It is a weak one while no number is attached.

State of the Art (as of 2026-09-05)

The named ingredient with no published means of production now has one. This page's claim is that the next increment of capability is bought after pre-training — from RL, generated training environments, and the harness the model is trained inside. Environments have been the weakest link in that list: every account read has assumed them and none has said where they come from at scale.

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments (2026-09-03) proposes manufacturing them out of a byproduct every lab running coding agents already holds in quantity. The stated asymmetry: agent trajectories have accumulated at scale while realistic executable environments remain scarce — and environments are what post-training actually consumes, because each can be re-queried into many verifiable tasks and returns execution feedback, whereas a trajectory is a single frozen demonstration. The method turns on the observation that a trajectory's tool-execution history exposes the structure and contents of the environment it ran in: replay the recorded file operations to restore each file to its pre-modification state, yielding a partial workspace, then have a completion agent supply the missing files and dependencies. Tasks are then scaled along breadth (mining directional dependency relations between environments to synthesise cross-workspace queries spanning multiple codebases) and depth (extending a single-turn query into a multi-round session driven by a user agent) (source).

Recorded with its central weakness stated: the abstract carries no figures at all. No benchmark, no baseline, no environment count, no success rate. Whether a model post-trained on reconstructed environments beats one trained on the trajectories they came from is the entire practical question and it is not addressed. Two structural doubts belong beside it — a completion agent that guesses a dependency wrong produces an environment that is executable and not the one the trajectory ran in, and environments are recoverable only where trajectories exist, so the available distribution is the distribution of tasks agents have already been asked to do, which is the opposite of the coverage argument the method is motivated by.

And the third ingredient — the harness — got its first controlled comparison against the second. WHALE: A Simple Recipe for Joint Harness-Weight Optimization (2026-08-31) alternates weight updates and harness search rather than scaling either alone, reporting +4.15–24.38 pp over weight-only, harness-only and Fast-Slow Training, and — the load-bearing finding — that either component can be the bottleneck depending on the domain. For this page that is a constraint on the scaling claim: post-training capability is not a single quantity that can be bought with more of one input. Scaling RL against a frozen harness runs into the harness; scaling the harness against frozen weights runs into the weights.

Neither paper cites the other, and neither cites Z.ai's claim. The composition is this wiki's.

State of the Art (as of 2026-08-21)

The vendor claim, with its evidence. Z.ai states GLM-5.3's base is GLM-5.2's, untouched, with every reported gain from scaled post-training — reported as RL on long-horizon environments. Vendor-stated, no harness published for any figure (source):

BenchmarkGLM-5.2GLM-5.3
Terminal-Bench 3.04.628.3
DeepSWE46.266.9
AutomationBench26.248.2
Terminal-Bench 3.0 (4.6 → 28.3) and DeepSWE 66.9 **match the figures captured
independently on 2026-08-14** and held on GLM-5.3
(source). The GLM-5.2 baselines
for DeepSWE and AutomationBench are new.

The research results, which arrived without reference to the vendor claim. Four papers in three days move a different component and hold the weights, or the base, fixed:

PaperWhat is scaledWhat is frozenHeadline
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393)RL inside three production harnessesthe harnesses' control flowSWE-bench Verified 64.0→70.4, 62.4→68.2, 57.2→66.6
Agent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528)RL through an endpoint proxythe harness owns the loopQwen3.5-9B 41.8→56.4 on SWE-bench Verified
SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197)the training environments themselves—+5.3 avg / +13.9 ACEBench-Agent at 30B
DarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545)the harness populationthe model~+17 avg over four benchmarks
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence (arXiv:2608.16590)runtime critics and recovery skillsthe base policyLIBERO-Pro 90.8%, RoboCasa 93.6%
2026-08-23 — the first third-party number, and it is smaller than the vendor's.
Artificial Analysis's weekly capture lists **GLM-5.3 (max) at Intelligence
Index 60** against GLM-5.2 (max) at 53 — a **7-point move on an
independent composite between two models Z.ai states share an unchanged base**
(source). GLM-5.3 was absent
from the 2026-08-16 capture, so this is its first appearance.

That is corroboration of the direction and a correction of the magnitude. Z.ai's own framing is a ~50% coding improvement and a six-fold Terminal-Bench move; a third party measuring a broad composite finds +7 points, which is real, is larger than the gap between most adjacent rows in that table, and is not the same claim. The honest reading is that post-training moved this model up roughly one tier — from GLM-5.2's neighbourhood into a band it shares with Kimi K3 (max) and Grok 4.6 (xhigh) — without a new pretrain.

It still does not supply the unit. An index delta is an outcome, not a quantity of post-training, so Open Problem 1 below is untouched by it.

And the ceiling nobody has published. StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089) reached GPT-5.6 Sol xhigh 95.3% on Terminal-Bench 2.1 with weights untouched, for under $38 of adaptation. No result in this cluster reports where the returns stop, or what a second doubling of post-training buys after the first.

Open Problems

  1. The cheapest lever on this axis has a yield set by a relationship, not by effort. Distillation is the least expensive way to move a model in post-training, and Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647) reports that how far on-policy distillation reaches is governed by the teacher/student origin relationship — same-origin pairs transfer across languages, horizons and domains; cross-origin pairs mostly fit the trained distribution — while combining teachers produces a mixture-dependent seesaw rather than their union, because routing cannot confine a teacher's influence. Nothing in the GLM-5.3 account says whether distillation is among the post-training this page is arguing about, so this is a constraint on the lever, not a claim about that model (source)

  2. What is the unit? "Scaled post-training" names a direction with no quantity. Pre-training scaling laws are stated in parameters, tokens and FLOPs; not one source here gives post-training compute, environment count or rollout count. Until one does, this is a claim about where effort went, not a law.

  3. Is the base actually untouched? The strongest form of Z.ai's claim is falsifiable and, as of today, testable: Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929) verifies weight lineage from checkpoints alone, data-free, at AUROC 1.0 on its benchmarks. If GLM-5.3's weights ship (staged for ~2026-08-28) the claim can be checked rather than believed. Nobody has done this, and the method has not been demonstrated at MoE frontier scale.

  4. Five knobs, none enumerated. Tang is reported to name five scaling knobs and an "XA-YB" MoE sparsity notation; nothing read defines either, so the framework cannot be applied by a reader.

  5. Does it survive independent measurement? Every GLM-5.3 figure here is vendor-stated with no harness published — the exact deficiency Eval Harness Configuration exists to flag, in the release that is the thesis's main evidence.

  6. Does capability transfer off the harness it was trained in? ClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798) measured a 1.48× spread from harness alone after training; LEGO-RL reports three separately-trained checkpoints and no cross-harness matrix. If post-training gains are harness-specific, "post-training scaling" and "harness scaling" are the same axis under two names.

  7. What happens to the open-weight bargain? If the base is the cheap half, a lab can publish weights, satisfy every open-weights commitment Open-Weights Policy Fight tracks, and retain the part that matters. Z.ai's staged release — base shared with GLM-5.2, post-trained weights held pending a safety evaluation — is the first instance shaped exactly that way, and no one has said whether that is the intent.

  8. Nothing here measures what post-training takes away. Every figure on this page is a top-1 score — Terminal-Bench 3.0 4.6 → 28.3, and the rest — which says how often the post-trained model gets it right and nothing about how much of the base model's reachable behaviour survived. Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351) (2026-08-27) reports that GRPO exhibits entropy collapse while Evolution Strategies does not, and that ES "improves Pass@1 while attaining higher Pass@K" — Pass@K being the measure of what remains reachable. If that holds, the dominant post-training method buys the headline number partly by narrowing the model, and this page's central claim would be resting on a quantity nobody in it has reported. Stated as an open problem, not a correction: the paper names no benchmark, no model and no absolute figure, and none of the vendor results here publish a Pass@K at all, so the two cannot yet be put side by side (source)

Key Papers

  • Scaling Properties of Same-Family On-Policy Distillation — scaling laws for same-family OPD; weak-to-strong students exceed their teachers, smaller teachers transfer better at matched score

  • Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation — 2026-09-08, and the third entry in twelve days on the same problem: what to do with a teacher signal that cannot be trusted as a target. On-Policy Reverse Distillation takes the hardest version — the teacher is genuinely weaker — and refuses to imitate it. It reads the teacher's policy shift relative to its own reference policy, on the student's rollouts, and amplifies only the component of the student's verifier-driven policy gradient already pointing that way. Because it rescales only verifier-supported updates, it preserves the stationary points of policy optimization, so a weak teacher cannot impose its ceiling by construction. Reported as higher performance with fewer student updates in successive model transfer and multi-teacher distillation, and a response-style analysis finds students closer to verifier-RL-only models than to their teachers. Read against the two entries below, the verifier is progressively taking authority from the teacher — bias to manage, then sign, now admission gate. No absolute figure, benchmark or model name appears in anything read, and none of the three cites either of the others (source)

  • NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness — 2026-09-08. The routing layer becomes the training corpus: each turn's predicted capability demand, selected service tier and subsequent interaction are recorded, admitted through structural validation and six-dimensional semantic evaluation, and replayed as examples that preserve interleaved reasoning, tool calls and harness context, with evaluation feedback setting the next training mixture. Macro-average across eleven benchmarks moves 58.94 → 64.87 at 4B and 65.60 → 69.04 at 9B, putting the post-trained 4B within 0.73 of the 9B base. The "recursive" in the title covers one closed loop — the paper calls itself "an initial prototype" — so the question the entry below raises about repeated self-generated signal is precisely the one it does not answer. No benchmark is named and no per-benchmark figure is published (source)

  • One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation — 2026-08-26. A review, with no new experiments, of On-Policy Self-Distillation: the teacher is the model itself, conditioned on privileged information the student will not have at test time, so it is no stronger, only better informed — which removes the cost of a second large model and is why the technique spread. Its argument is that the asymmetry producing the signal also biases it, and that the field's dominant failure, collapse (the progressive narrowing of the reasoning paths a model can produce), is one symptom governed by three levers: where the signal is applied, what the teacher is shown, and when the guidance decays. Collapse is not specific to OPSD; privileged information aggravates it. Scope is restricted to mathematical reasoning (source)

  • FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience — 2026-09-03, and read as the pair to the entry above. FlowBalance keeps dense self-guidance but lets a verifier decide its sign: guidance is retained on positive-advantage trajectories, reversed on negative-advantage ones, and switched off when the rollout group expresses no outcome preference — realised through profiled trajectory balance, with no separate token-level imitation loss. Reported: better average performance than FlowRL on Qwen3-4B and Qwen3-8B, improved training speed and stability, avoidance of direct OPSD's response-length collapse, and higher correct-strategy diversity on a controlled AIME24 diagnostic. No absolute figures appear in anything read, and the two papers do not cite each other — the pairing is this wiki's. It moves exactly one of the review's three levers, which is what makes it a test of the review's claim rather than an illustration of it. Reporting diversity at all is the rarer thing: accuracy on a held-out set cannot see collapse (source)

  • SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers — 2026-09-03. The control this field usually omits: looped-Transformer comparisons are normally run at fixed model size, which hands the looped model extra FLOPs. SMELT matches per-token FLOPs, non-embedding parameters and KV cache, and the advantage survives — 6.8–18.0% of training FLOPs saved on the compute-optimal frontier, fitted as a separate Chinchilla-style law across four sizes to 54B non-embedding parameters. Two findings cut against how scaling results are usually read: the downstream gain exceeds what validation loss predicts (largest on Code), and it grows with sample length and in-context examples, so the architectures are not comparable at a single sequence length (source).

  • LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393) — RL inside three named production harnesses; supplies the harness identity Agent Lightning omitted.

  • Agent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528) — names harnessed agentic RL and reports four other frameworks adopting the architecture.

  • SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197) — makes the training environment a learned component.

  • Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (arXiv:2608.17253) — removes ground-truth labels from RL via peer reward across a diverse cohort.

  • DarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545) — harness evolution with the model frozen.

  • StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089) — the cost figure the cluster otherwise lacks: $15 vs $574.68.

  • Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351) — the only entry here that asks what post-training costs rather than what it buys: GRPO's entropy collapse against Evolution Strategies' higher Pass@K.

  • Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929) — the instrument that could test claim 2 above.

  • Rufus-Air: An Open LLM Post-Training Recipe — Rufus-Air: An Open LLM Post-Training Recipe, 2026-09-24. Eight serial stages on GLM-4.5-Air-Base, no new human annotation and no in-house distillation teacher, ordering principle reward reliability — and no benchmark figure of any kind (source)

Referenced by

Adversarial DistillationAgentGrad: Intervention-guided Prompt Optimization for Multi Agent SystemsAgentic Reinforcement LearningAgents (LLM Agents)ApodexApodex 1.1: Scaling Agentic Intelligence for Complex Work (arXiv:2608.23283)August 2026 — Monthly DigestBeneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMsBeyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311)Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (arXiv:2608.17253)EnvHarness: Awakening Static Worlds for Agent Learning (arXiv:2608.19880)Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647)FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning ExperienceGLM-5.3Learning to Solve Hard Problems in RL for LLMs by Never Giving UpLEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393)Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts (arXiv:2608.20061)Liquid AINegative Self-Distillation: Learning to Reason by Avoiding FlawsNeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing HarnessOne Symptom, Three Levers: A Critical Review of On-Policy Self-DistillationOpen-Weights Policy FightPaul ChristianoRethinking Critic Learning in PPO: Understanding and Mitigating Value FlatteningRewardVerse: Rubric-Guided Policy Optimization for Video Reward ModelingRufus-Air: An Open LLM Post-Training RecipeScaling Properties of Same-Family On-Policy DistillationSMELT: Scaling Laws for Compute-Matched MoE Looped TransformersSPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197)Terminal-Universe: Turning Agent Trajectories into Scalable Terminal EnvironmentsUnderstanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351)Weekly Synthesis — W34 (2026-08-17 → 2026-08-23)Weekly Synthesis — W37 (2026-09-07 → 2026-09-13)When EOS Tokens Disagree: Understanding Length Inflation in On-Policy DistillationWikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill EvolutionZ.ai

Sources