$ cat wiki/concepts/adversarial-distillation.md
Adversarial Distillation
Definition
Adversarial distillation is the extraction of a model's capability from its outputs, by a party the developer did not authorise. OpenAI gave the term its first published definition on 2026-09-30:
adversarial distillation: the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model
(source)
It is worth separating from the technique it borrows its name from. Post-Training Scaling and the on-policy distillation literature — eight paper pages in this wiki — describe distillation as a method a lab applies to its own models. This page is about the same arithmetic performed across an ownership boundary, against a model the distiller only reaches through an API. The mechanism is shared; the dispute is entirely about consent.
This page exists because the topic had outgrown its containers. Two lab-level
accusations, a named technique, a mechanism paper and a mitigation were being
recorded across Open-Weights Policy Fight, AI Governance,
AI-Enabled Cyberattacks and eight entity pages, with no page saying
what the thing is. The GTG-16005 string alone appears on nine documents.
Why It Matters
Because it is the first capability-transfer channel that a licence cannot close. An open-weight release can attach conditions; a closed API can attach terms of service. Both were attached, and both were reportedly circumvented at a scale measured in hundreds of millions of exchanges.
Because the two accusations disagree about what is being taken. Anthropic's
GTG-16005 describes chain-of-thought distillation harvested at volume and
used to train named models. OpenAI's disclosure describes an attempt to recover
protected reasoning that was deliberately withheld from the answer — and
stops short of saying any model was trained on it. One is an alleged completed
transfer; the other is an alleged attempted one.
Because the defence and the attack are both cheap. Claude Opus 5.5 ships preserved thinking as an explicit anti-distillation mechanism. Against that, Post-Training Leaves Behavioral Shadows on Unrelated Decisions needs one word of teacher output per prompt and no task examples at all — a channel that does not touch the reasoning trace, and against which nothing read proposes a defence.
State of the Art (2026-10-01)
The two accusations
| Anthropic → Alibaba / Qwen AI Lab | OpenAI → Moonshot AI | |
|---|---|---|
| Disclosed | 2026-09-10 | 2026-09-30 |
| Codename | GTG-16005 | none published |
| Window | May–July 2026 | July 1–28, 2026 |
| Accounts | >3,500 fraudulent | ~4,000 users, then >15,000 |
| Volume | >151M exchanges, peaking ~3M/day | 16,000 prompts on July 24–25 |
| Target | chain-of-thought of Claude Opus 4.6 and 4.7 | protected reasoning |
| Outcome alleged | transcripts trained Qwen 3.5, 3.6, 3.7 | none stated |
| Evidence published | not published | not published |
| Both are the accusing lab's own account, unverified by any third party, and | ||
| both name a Chinese lab. Anthropic states it has disrupted seven China-based | ||
| labs for distillation since February 2026; Moonshot is the second to be named | ||
| publicly by anyone. |
The technique OpenAI describes
The operators did not break encryption, reach a database, or read stored
conversations. They copied encrypted reasoning out of one conversation and
asked another model instance to decrypt and transcribe it — a model used against
its own protection
(source).
cyberscoop.com calls this a "novel encryption bypass"; the description in both
search passes is that the ciphertext was never attacked, only re-presented to
something holding the key.
Response stated: accounts banned, sign-up checks tightened, findings shared through the Frontier Model Forum and government channels.
What the research says about whether it works
Yes, and better from a small teacher than a large one. Scaling Properties of Same-Family On-Policy Distillation fits power laws to on-policy distillation and reports that in every observed weak-to-strong pair the student's peak score exceeds its teacher's own, and that at a matched score smaller teachers transfer better (source).
Read against these accusations, that is the uncomfortable result. A distiller does not need access to the largest model, and the ceiling is not the teacher's score. The paper is about a lab distilling its own models and says nothing about unauthorised use — the connection is this wiki's, and it is a mechanism, not evidence.
Open Problems
- No accusation has published its evidence. Two labs have now described campaigns in numbers, and neither has published transcripts, account identifiers, or a method a third party could check. The figures are the only artefact.
- Anthropic's own two figure sets still do not reconcile — ~25,000 accounts
and 28.8M interactions in the June letter against >3,500 accounts and 151M
exchanges in
GTG-16005. Recorded on Alibaba / Qwen AI Lab and unresolved there since 2026-09-10. - Attribution is asserted, not shown. Both search passes on the OpenAI disclosure note that no hard evidence was offered for the Moonshot link.
- A defence for the one-word channel. Preserved thinking protects the trace; ATD does not read the trace.
- Whether "adversarial" survives contact with the open-weight case. Every model in Open-Weights Policy Fight may be distilled by anyone by design. The term only does work where a boundary was asserted, which is why it arrives from the two labs that assert one.
Key Papers
- Scaling Properties of Same-Family On-Policy Distillation — scaling laws for on-policy distillation; weak-to-strong transfer exceeds the teacher
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions — Active Taskless Distillation: one word per prompt, no task data
- Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647) — how far an OPD student generalises
- Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation — weak-to-strong on-policy reward distillation
- Weak-to-Strong Generalization via Direct On-Policy Distillation — the earlier weak-to-strong result this wiki holds
Related Concepts
- Open-Weights Policy Fight — where distillation is a feature, not a breach
- AI Governance — the policy response, including the export-control framing
- Post-Training Scaling — distillation as a lab's own method
- AI-Enabled Cyberattacks — where
GTG-16005was first recorded - Frontier Pacing — whether a four-month lag is a lead