$ cat wiki/papers/2026/2609.35259-distillation-dynamics.md
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
TL;DR
A controlled study that varies rollout policy, token-level KL direction and learning rate independently, and finds that rollout policy is not the variable that matters most. KL direction shapes performance and coverage; learning rate governs forgetting and update sparsity.
Authors & Org
unknown. arxiv.org answers EGRESS_BLOCKED from this pipeline, so the paper was
not read and the abstract is the only text available, via the HuggingFace Daily
Papers snapshots of 2026-10-03 and 2026-10-04
(source).
No author list or affiliation appears in anything read and none is guessed.
Model families used: Llama3 and Qwen2.5 — both open-weight.
Method
A strong-to-weak distillation setting, with three factors varied independently rather than jointly:
- rollout policy — teacher-generated vs student-generated, and a continuous student–teacher spectrum between them
- token-level KL direction — forward vs reverse
- learning rate
across the Llama3 and Qwen2.5 families and reasoning tasks spanning scientific, medical and arithmetic domains. KL gradients are analysed directly.
The design is the contribution. The abstract's stated complaint about prior work is that comparisons between supervised fine-tuning and reinforcement learning "vary many factors simultaneously, making the contribution of rollout policy difficult to isolate" — so the on-policy advantage reported in the literature may be an artefact of confounded comparisons.
Robustness checks reported: removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains.
Results
From the abstract (source). No numeric figure of any kind appears in anything read — no accuracy, no forgetting rate, no sparsity measure. Every result below is directional.
- Rollout policy "does not necessarily play a central role."
- Token-level KL direction "more clearly shapes task performance and output coverage."
- Learning rate "governs forgetting and update sparsity" — the two effects most often attributed to on-policy training.
- Forward KL is "remarkably robust to rollout policy", stable and strong regardless of where rollouts come from.
- Reverse KL is "substantially more sensitive and favours student-generated rollouts."
- On-policy data does help generalisation to harder variants of the Countdown arithmetic task, under both KL directions — but "this advantage does not reliably persist after subsequent RLVR."
The stated conclusion: results "challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters."
The forward/reverse KL split is what makes the headline coherent rather than merely contrarian. If forward KL is insensitive to rollout policy and reverse KL is not, then "is on-policy better?" has no answer independent of the objective — which is why confounded comparisons could report either result honestly.
Significance
This is a consensus-overturning result in the sense interests.md tracks: the
claim that on-policy rollouts reduce catastrophic forgetting, sparsify updates and
improve generalisation is widely held, and this paper reassigns each of those three
effects to a different knob. Forgetting and sparsity go to the learning rate;
performance and coverage go to KL direction.
For Agentic Reinforcement Learning and the post-training material on Test-Time Compute (Inference-Time Compute Scaling), the practical reading is that an on-policy vs off-policy decision is not separable from the KL objective it is made under, and a paper reporting one without the other is reporting an underdetermined experiment.
The Countdown finding deserves its own weight: on-policy data does buy generalisation to harder variants, and then that advantage does not survive RLVR. A gain that disappears under the next training stage is a different kind of claim from one that does not exist, and the paper reports it as the former.
Open Questions
- No numbers. Every finding is directional. Effect sizes, and therefore whether any of this is practically large, cannot be assessed from what was read.
- Llama3 and Qwen2.5 only, both open-weight and neither frontier-class. Whether the KL-direction asymmetry holds at frontier scale is untested.
- Strong-to-weak distillation only. The result is not claimed for same-size or weak-to-strong settings, and this wiki holds Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation in the adjacent direction.
- The RLVR stage that erases the Countdown advantage is not characterised, so it is unclear whether this is specific to RLVR or general to any subsequent RL.
Cite
arXiv:2609.35259, published 2026-09-28. Surfaced via HuggingFace Daily Papers in the 2026-10-03 and 2026-10-04 snapshots, at 153 upvotes in today's — a popularity signal from that community and nothing more (source).
Captured 2026-10-04, +6 days. It appeared in the 2026-10-03 snapshot, which no run consumed — there was no pipeline run on 2026-10-03.