$ cat wiki/papers/2026/2609.32722-opd-scaling-properties.md
Scaling Properties of Same-Family On-Policy Distillation
TL;DR
Power laws for on-policy distillation, and the headline result is that a small
RL expert can lift a much larger student past its own score. Across
weak-to-strong, same-base and strong-to-weak teacher–student pairs, early OPD
training shows a regular useful-transfer regime in which held-out accuracy
(the gold score, G) rises approximately linearly in the square root of
token-level reverse KL divergence from the student initialization. In every
observed weak-to-strong pair the student's peak gold score exceeds its teacher's
own, and at a matched gold score smaller teachers transfer better — so a
teacher's score alone does not define its supervision value
(source).
194 upvotes, the top entry in today's snapshot — a popularity signal from the HuggingFace community and nothing more.
Authors & Org
Not stated. The HuggingFace Daily snapshot carries no author block and
arxiv.org answers EGRESS_BLOCKED from this run's sandbox. Recorded as unknown
rather than inferred.
Method
The question is how much of the reasoning capability that RL induces transfers across model scales, and how fast. Three teacher–student configurations:
| Setup | Teacher vs student |
|---|---|
| weak-to-strong | teacher smaller than student |
| same-base | same base model |
| strong-to-weak | teacher larger than student |
| The measurement device is the useful-transfer regime: early in OPD training, | |
G is approximately linear in d = sqrt(KL(π_θ ‖ π_ref)) — reverse KL from the | |
| student's initialization, not from the teacher. Two quantities are then fitted | |
| as power laws against student parameter count, teacher parameter count and teacher | |
| gold score: |
G_peak— the student's best held-out accuracy- the slope of the useful-transfer regime — how fast it gets there
Two variants are also studied: bootstrapping weak-to-strong OPD, and varying the degree of on-policy supervision.
Results
| Claim | Finding |
|---|---|
| Transfer regime | Uniform across all observed setups; G linear in d |
| Weak-to-strong | Student peak exceeds the teacher's own score in every observed pair |
| Teacher scale | G_peak improves with teacher scale only up to roughly the student's scale |
| Teacher score | At matched gold score, smaller teachers transfer better |
| **No absolute accuracy figure, model family, or parameter count is quoted in the | |
| abstract this snapshot carries.** The results above are the shape of the scaling | |
| laws, not points on them. That limit is the page's, not the paper's — the fitted | |
| constants are presumably in the paper body, which this pipeline cannot read. |
Significance
It is a capability-transfer result read in a week when two labs accused others
of performing exactly this transfer without permission. OpenAI disclosed the
Moonshot-linked campaign on 2026-09-30, the day before this snapshot; Anthropic
disclosed GTG-16005 on 2026-09-10. This paper says the arithmetic works, that a
distiller does not need the largest teacher, and that the student's ceiling is not
the teacher's score. See Adversarial Distillation.
The paper is about a lab distilling within its own model family and says nothing about unauthorised use. The connection is this wiki's; it supplies a mechanism, not evidence, and the distinction is the same one drawn on Post-Training Leaves Behavioral Shadows on Unrelated Decisions yesterday.
Read against Post-Training Scaling, it cuts the other way from Z.ai's claim. Post-training scaling argues that capability keeps coming from the RL stage; this says the product of that stage is cheaply movable, and most cheaply by the smallest thing that has it.
The sqrt reverse-KL parameterisation is the reusable part. It gives OPD an
x-axis that is measurable from the student alone, without reference to the teacher
— which is what makes the weak-to-strong comparison meaningful at all.
Open Questions
- Where the laws break. "Only up to roughly the student's scale" is a turning point; the abstract does not say what happens past it.
- Why smaller teachers transfer better at matched score. Stated as a finding, with no mechanism offered in the abstract.
- Whether this survives the family boundary. "Same-family" is in the title; ATD works across families and generations, and nothing here speaks to that.
- Cost. No compute figure for reaching
G_peakby OPD against running RL on the student directly — which is the comparison a lab would actually make.
Cite
arXiv 2609.32722, published 2026-09-26, captured from HuggingFace Daily Papers 2026-10-01 at 194 upvotes (source).