AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/adversarial-distillation.md

Adversarial Distillation

conceptupdated 2026-10-01created 2026-10-01

Definition

Adversarial distillation is the extraction of a model's capability from its outputs, by a party the developer did not authorise. OpenAI gave the term its first published definition on 2026-09-30:

adversarial distillation: the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model

(source)

It is worth separating from the technique it borrows its name from. Post-Training Scaling and the on-policy distillation literature — eight paper pages in this wiki — describe distillation as a method a lab applies to its own models. This page is about the same arithmetic performed across an ownership boundary, against a model the distiller only reaches through an API. The mechanism is shared; the dispute is entirely about consent.

This page exists because the topic had outgrown its containers. Two lab-level accusations, a named technique, a mechanism paper and a mitigation were being recorded across Open-Weights Policy Fight, AI Governance, AI-Enabled Cyberattacks and eight entity pages, with no page saying what the thing is. The GTG-16005 string alone appears on nine documents.

Why It Matters

Because it is the first capability-transfer channel that a licence cannot close. An open-weight release can attach conditions; a closed API can attach terms of service. Both were attached, and both were reportedly circumvented at a scale measured in hundreds of millions of exchanges.

Because the two accusations disagree about what is being taken. Anthropic's GTG-16005 describes chain-of-thought distillation harvested at volume and used to train named models. OpenAI's disclosure describes an attempt to recover protected reasoning that was deliberately withheld from the answer — and stops short of saying any model was trained on it. One is an alleged completed transfer; the other is an alleged attempted one.

Because the defence and the attack are both cheap. Claude Opus 5.5 ships preserved thinking as an explicit anti-distillation mechanism. Against that, Post-Training Leaves Behavioral Shadows on Unrelated Decisions needs one word of teacher output per prompt and no task examples at all — a channel that does not touch the reasoning trace, and against which nothing read proposes a defence.

State of the Art (2026-10-01)

The two accusations

Anthropic → Alibaba / Qwen AI LabOpenAI → Moonshot AI
Disclosed2026-09-102026-09-30
CodenameGTG-16005none published
WindowMay–July 2026July 1–28, 2026
Accounts>3,500 fraudulent~4,000 users, then >15,000
Volume>151M exchanges, peaking ~3M/day16,000 prompts on July 24–25
Targetchain-of-thought of Claude Opus 4.6 and 4.7protected reasoning
Outcome allegedtranscripts trained Qwen 3.5, 3.6, 3.7none stated
Evidence publishednot publishednot published
Both are the accusing lab's own account, unverified by any third party, and
both name a Chinese lab. Anthropic states it has disrupted seven China-based
labs for distillation since February 2026; Moonshot is the second to be named
publicly by anyone.

The technique OpenAI describes

The operators did not break encryption, reach a database, or read stored conversations. They copied encrypted reasoning out of one conversation and asked another model instance to decrypt and transcribe it — a model used against its own protection (source). cyberscoop.com calls this a "novel encryption bypass"; the description in both search passes is that the ciphertext was never attacked, only re-presented to something holding the key.

Response stated: accounts banned, sign-up checks tightened, findings shared through the Frontier Model Forum and government channels.

What the research says about whether it works

Yes, and better from a small teacher than a large one. Scaling Properties of Same-Family On-Policy Distillation fits power laws to on-policy distillation and reports that in every observed weak-to-strong pair the student's peak score exceeds its teacher's own, and that at a matched score smaller teachers transfer better (source).

Read against these accusations, that is the uncomfortable result. A distiller does not need access to the largest model, and the ceiling is not the teacher's score. The paper is about a lab distilling its own models and says nothing about unauthorised use — the connection is this wiki's, and it is a mechanism, not evidence.

Open Problems

  • No accusation has published its evidence. Two labs have now described campaigns in numbers, and neither has published transcripts, account identifiers, or a method a third party could check. The figures are the only artefact.
  • Anthropic's own two figure sets still do not reconcile — ~25,000 accounts and 28.8M interactions in the June letter against >3,500 accounts and 151M exchanges in GTG-16005. Recorded on Alibaba / Qwen AI Lab and unresolved there since 2026-09-10.
  • Attribution is asserted, not shown. Both search passes on the OpenAI disclosure note that no hard evidence was offered for the Moonshot link.
  • A defence for the one-word channel. Preserved thinking protects the trace; ATD does not read the trace.
  • Whether "adversarial" survives contact with the open-weight case. Every model in Open-Weights Policy Fight may be distilled by anyone by design. The term only does work where a boundary was asserted, which is why it arrives from the two labs that assert one.

Key Papers

Referenced by

Sources