AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/open-weights-policy.md

Open-Weights Policy Fight

Definition

The 2026 policy dispute over whether openly released model weights should be restricted, and on what grounds. The question is no longer academic: it now has an industry coalition on one side, three frontier labs and a Treasury sanctions threat on the other, and a live security incident that both camps cite as evidence.

Why It Matters

  • The industry split is now formal. Until July 2026 the disagreement was rhetorical. The launch of the Open Secure AI Alliance put roughly forty companies behind open weights as a security position, and made the absence of the frontier labs a matter of public record.
  • It is being argued from incident evidence, not preference. Both sides now point at the same July 2026 agent intrusion — one to argue open weights contained it, the other to argue agentic capability needs gating. → AI-Enabled Cyberattacks
  • The distillation dispute is the geopolitical face of the same fight. Restricting open weights and sanctioning Chinese labs for distillation are the same lever pointed in two directions. → AI Governance

State of the Art (2026-10-04)

A European lab shipped Apache-2.0 weights whose selling point is how little of the model runs at a time. Aleph Alpha released Kolibri-1 on 2026-10-03 — 78.1B total parameters, 3.46B active per token, 384 experts with 6 active, 50 layers, Apache 2.0, full weights on Hugging Face (source). Read first-party; aleph-alpha.com answered normally.

3.46B active is the number that matters to this page. Every open-weight release tracked here trades off licence, capability and what it costs to run, and the third term is set by active parameters rather than by total. Aleph Alpha's claim is that the model

matches models with up to four times its active parameter count

across math, coding, grounding and long-context tasks — and its own table backs that unevenly: AIME 2025 96.9 and GPQA Diamond 84.3 against Nemotron 3 Super 120B's 91.7 and 78.0, but HumanEval+ 92.7 against Nemotron's 94.7, and BFCL v4 61.4 against 61.0, which is a tie. So the efficiency claim is carried by maths and science and not by code or tool use. The training mix is consistent with that: ~14% code against ~62% English and 21.3% German.

The sovereignty argument here is measurable rather than rhetorical, which is unusual on this page. Of a 20-trillion-token three-stage pre-training run, 4.3 trillion tokens are German — the highest disclosed non-English share of any model this wiki holds. Where Mistral AI's European-sovereignty case is made in prose and funding announcements, this one is made in a data-mix percentage that can be checked against the benchmark split.

Two things are not stated, and both bound the release: no price and no hosted endpoint (so the distribution route is download the weights, and Pricing is unknown on the model page), and no comparison against any frontier closed model. The comparison set is two ~120B open models.

The native context window is disputed and the dispute is not resolved. The first-party post states 16,384 tokens native, extended to 1,048,576; secondary coverage states 262,144 native. The first-party figure is the one carried, per the CLAUDE.md priority rule, and the disagreement is recorded on Kolibri-1's ## Conflicting Reports — a 16× gap in native length changes what the extrapolation to 1M is doing, so it is not a rounding difference.

State of the Art (2026-10-01)

The safeguards on an open-weight frontier model were removed for $4,400 by a team that had never done it before. Anthropic's Frontier Red Team published the figure on 2026-09-29, read first-party on this run (source).

Engagement with malicious cyber-attack orders in a simulated environment:

ConditionGLM-5.3Claude
Bare order0%0%
False cover story64%0%
Prefilled reasoning92%0%
Abliterated100%n/a
Verbatim: "Abliterating the model took our team—which had never previously attempted this task—about 2,200 GPU hours at a computation cost of roughly $4,400." Refusal rates afterwards, against an original above 90%:
JailbreakBench 3% · HarmBench 2% · StrongREJECT 12%.

This page has argued the release-conditions question in the abstract for three months; the figure is now concrete and it is small. $4,400 is not a state-actor budget. And the 64% and 92% rows matter more than the 100%: a false cover story and prefilled reasoning tokens require no GPUs at all — the safeguards were already permeable before anyone touched the weights.

The asymmetry Anthropic states about itself is the load-bearing sentence, and it is also a competitor describing its own advantage:

Claude models are released with cyber safeguards, and versions with reduced safeguards are limited to vetted users. Anyone can download and use GLM-5.3.

Read against this page's existing record, that is the open-weights bargain stated at its least favourable. Every argument here for permissive release — reproducibility, sovereign capability, price discipline, the audit surface — is made about weights anyone can modify, and abliteration is the cheapest possible modification.

What it does not establish. The post publishes no comparison of abliteration cost across open-weight models, so $4,400 is one data point on one model, not a curve. Nothing read measures whether an abliterated GLM-5.3 retains the 12% ExploitBench rate — the capability figures and the safeguard figures are reported separately, and this page does not combine them.

CAISI's four-month lag figure now has to carry two things. NIST's Center for AI Standards and Innovation calls GLM-5.3 "the most cyber-capable open-weight model released to date", four months behind the US frontier. A four-month lag with safeguards and a four-month lag without them are not the same policy object, and Frontier Pacing has been measuring the first.

State of the Art (2026-09-28)

A derivative narrowed an Apache-2.0 licence, and the party that did it is not a lab (2026-09-26)

Every narrowing this page records was done by the lab that trained the model — Alibaba / Qwen AI Lab on Qwen-Image-2.1, Z.ai's revenue-gated security-review clause on GLM-5.3, the 2026-08-17 "one vendor, two licences" entry. This one was done downstream, by someone who did not train anything.

UkisAI's Swift 1.5 is an adaptation of Qwen 3.8 27B, whose base is "Copyright 2026 Alibaba Cloud, Apache License 2.0". The adapted weights are released under the Swift Open License v1.0: gated access, free for personal, research, educational, evaluation and commercial use up to US$1,000,000 of annual revenue, and a separate Swift Enterprise License required above that threshold (source).

Why it belongs on this page. Apache-2.0 permits exactly this — a permissive licence is permissive about being built on by a stricter one, and nothing here is a violation. What it shows is that the licence of a weight set is not a property of the model, it is a property of the last party to touch it, and this page has been arguing about labs' licences as though they governed what reaches a user. They govern the first hop. A revenue gate reappearing one hop downstream, on a base chosen because it was Apache-2.0, is the mechanism by which an open ecosystem produces gated artefacts without any lab changing its mind.

The revenue threshold is the same shape as GLM-5.3's and the trigger is different. GLM-5.3's clause gates on revenue to require a security review; this one gates on revenue to require a paid licence. Same instrument, one used for safety conditioning and one for monetisation — which is the distinction this page's Open Problems ask for and rarely gets two clean examples of.

Reported results, and the reason there is no model page. Two token-reduction figures at different stated settings: −58.5% thinking tokens at +0.35% quality with a 9.18× speed-up "on several tasks", and separately ≈−29% fewer thinking tokens at low reasoning effort while scoring above the base. Evaluation covers GPQA-Diamond (198 questions), IFBench (300 prompts) and AIME 2026 (30 problems), over the merged checkpoint and three INT4 exports. No model page is created, for the reason Qwen 3.8 27B's series-mate Hunyuan-A13B was refused one on 2026-09-27: huggingface.co is blocked from this run, every figure is a search-snippet of a model card, the context window is unread, and the release date is bounded (2026-09-26 from the r/LocalLLaMA post) rather than stated. A ## Spec table cannot assert a Released value from a relative "2 days ago". The licence fact needs no spec table, and it is the part that matters here.

State of the Art (2026-09-27)

Two more "open" releases, and the word means two different things in them.

Ling-3.0-tiny — Ant Group (inclusionAI / AntLing), released 2026-08-06, captured +52 days — is MIT, with weights in BF16, FP8, INT4 and GGUF. That is the second MIT release from this lab in two days of captures, after Ming-Image-0.1-Design, and MIT remains the most permissive licence this wiki holds on any Chinese release in any modality (source).

Hunyuan-A13B Technical Report — Tencent — describes an "open-source" 80B/13B MoE. Search passes give its licence as a custom Tencent licence whose territory excludes the European Union, the United Kingdom and South Korea (source). Not read first-party, and if confirmed it is a geographic carve-out rather than a field-of-use restriction, which is a category this page has not previously held: Qwen-Image-2.1's Qwen Research License restricts commercial use; a territory exclusion restricts who may use it at all.

So the September picture, in one table, from releases captured in the last seven days:

ReleaseDeveloperStated termsRestriction type
Ling-3.0-tinyAnt GroupMITnone
Ming-Image-0.1-DesignAnt GroupMITnone
Qwen-Image-2.1AlibabaQwen Research Licensenon-commercial
Hunyuan-A13BTencentcustom Tencent licence (unconfirmed)territorial
The reading this page will not make: that Chinese labs are diverging on licensing as a
matter of policy. Four releases from three labs inside one week, with one licence unread
and one release fifty-two days old, is not a trend — it is the observation that **"open weights" is being used for MIT and for a territory-excluded custom licence in the same
month**, and that this wiki can only tell them apart when someone reads the file.

A fact about attention rather than about licensing, and it belongs here because this page's premise is that permissive terms enable downstream work: Ternary Bonsai 2 27B exists because PrismML rebuilt a Qwen derivative under Apache 2.0. Ling-3.0-tiny has been MIT and downloadable for 52 days and nothing in this wiki has been built on it — nor, so far as anything read shows, anywhere else. A permissive licence is necessary and demonstrably not sufficient.

State of the Art (2026-09-26)

Two Chinese image models, two days apart, in opposite licensing directions (2026-09-20 / 2026-09-22)

Qwen-Image-2.1 shipped 2026-09-20 under the Qwen Research License Agreement, non-commercial — the first Alibaba / Qwen AI Lab release this wiki has recorded under terms that are not permissive, and a line this page had treated as reliably open.

Ming-Image-0.1-Design shipped 2026-09-22 under MIT — the most permissive licence this wiki holds on any Chinese image model — from Ant Group (inclusionAI / AntLing), publishing as inclusionAI. Two 6B models plus two open-source Agent Skills (source).

The pairing is the point, and it argues against reading licence policy as a national posture. Two Chinese labs, two days apart, in the same modality, moved in opposite directions. Whatever explains Alibaba's change is not a jurisdictional constraint, or Ant Group would be under it too.

What MIT makes possible here is concrete rather than notional. Ternary Bonsai 2 27B is this wiki's worked example: PrismML rebuilt a Qwen3.8-27B derivative under Apache 2.0 on 2026-09-17. That is not possible under the terms Qwen-Image-2.1 now carries, and is possible under MIT.

No stated reason for either licence choice appears in anything read — the absence was recorded on the Qwen page on 2026-09-21 and is still absent, and Ant Group published no rationale either.

Not established: no first-party document was read for the Ming-Image release — no model card, README or licence file; huggingface.co answers connect_rejected under standing policy and github.com was not reached. The MIT finding rests on two search passes, not on a licence file. Whether either licence reaches generated outputs as well as weights is unaddressed in both cases.

The gap this page argues about was put to Congress as five numbers, and this wiki can check every one of them (2026-09-21)

Nathan Lambert published his prepared congressional remarks on open-weight models, given as a briefing to Congressional members and staff on the state of open-weight models in the lens of U.S.-China competition, expanded into a public "state of the union on open models" (source).

The figures, quoted under the column heading they sit beneath, Artificial Analysis Intelligence Index, as read 2026-09-14:

ModelCreatorArtificial Analysis Intelligence Index
GLM-5.3Z.ai45
Kimi K3Moonshot AI44
GLM-5.3-FlashZ.ai42
InklingThinking Machines Lab26
Nemotron 3 UltraNVIDIA23
**Every one of the five matches this wiki's own capture of the same leaderboard,
cell for cell** (source),
one day earlier. That is the first time a third party's headline figures have
been checkable against a snapshot this repo already held, and they check out.
The two rows with effort tiers match at max.

Two further claims, each on one pass and neither checkable here: the top American open models were released in June and July 2026 and are updated less frequently than their Chinese counterparts, and Chinese labs release models with higher scores 2–6 months before American companies.

The index was revised between this page's older readings and this one, and comparing across them is invalid. This page still carries Inkling at 41 against Claude Opus 5 (max) at 61, read 2026-08-02 (source). By 2026-09-13 the top of the same board reads 53 and Inkling reads 26; Kimi K3 (max) went 57 → 44 over the same seven weeks. A model losing 15 points while the ceiling loses 8 is a rescale, not a regression, and neither the column heading nor the snapshot's own header text changed to say so. Both readings stay on this page with their dates, and no figure from one is set against a figure from the other.

What this does not settle. No first-party read — interconnects.ai and aiweekly.co both answer EGRESS_BLOCKED — so the framing above is a search extract with a pass count. The hearing, committee and date are not named in anything read, nor is whether a transcript exists. No adoption or download figure appears in any pass, though the essay is described as covering adoption patterns, and no policy recommendation is recorded: what Lambert asked Congress to do is not in anything read. Nothing outside the five-model table was recoverable.

A lab that has only ever widened its licences narrowed one, and did it without saying why (2026-09-20)

Alibaba released Qwen-Image-2.1 under the Qwen RESEARCH LICENSE AGREEMENT, dated September 20, 2026, licensed by Hangzhou Tongyi Laboratory Technology Co., Ltd. The grant is "FOR NON-COMMERCIAL PURPOSES ONLY"; "Non-Commercial" is defined in the document as "for research or evaluation purposes only"; commercial use requires a separately negotiated licence obtained at model-business@notice.qwencloud.com (LICENSE) (source).

Two passes state this is a change from the earlier Qwen-Image line, which shipped under Apache 2.0, with one adding that the Apache-licensed 2512 and Edit-2511 models are unaffected.

This is the first narrowing of terms by an incumbent open-weights lab recorded on this page. Every prior licence entry here has run the other way or held level — Alibaba's own Qwen3.8 opened on 2026-08-12, Qwen-Drive shipped Apache 2.0 across code, weights and demo data on 09-07, Thinking Machines shipped Apache 2.0 at frontier scale, Z.ai's GLM-5.2 base is MIT and downloadable. The two restrictions previously recorded were novel terms on new lines — MiniMax's territory exclusion (2026-08-03) and GLM-5.3's staged flagship hold-back — not a line that was permissive becoming less so.

It is the sharpest available test of the question two sections below, "Does open weights still transfer the capability?" That question asks whether a lab can satisfy its open-weights commitments while retaining the part that produced the capability. This is a blunter instrument than the post-training split Z.ai's CEO describes: the weights are published in full and the right to use them commercially is withheld outright. A downloadable checkpoint that cannot be deployed is open by every measure this page's trackers use — Mozilla's download and derivative counts included — and closed to every commercial user.

The concrete cost is on this wiki already. On 2026-09-17 this page recorded PrismML publishing Ternary Bonsai 2 27B, a ternary quantization of Qwen 3.8 27B, under Apache 2.0 — which it could do because Alibaba's licence permitted it. The same act on Qwen-Image-2.1 would require a negotiated licence from Hangzhou Tongyi. The derivative ecosystem this page measures as the benefit of open weights is exactly what a non-commercial clause removes.

What is not established, and it is the part that would settle the reading: no stated reason for the change appears in anything read, and whether the research licence reaches generated outputs as well as weights is not addressed in the clauses read. Nothing read says whether this is specific to the image line or a direction for Qwen generally — the language line's next release is the test.

The quantity this page argues about has been measured twice a year by someone else, and this wiki had captured neither reading (2026-09-15)

Mozilla published The State of Open Source AI v1.1 on 2026-09-15, with data current to 2026-09-01. It is the second edition; v1.0 went out 2026-07-14. A search of this entire wiki for "Mozilla" returned nothing before today — so a recurring, dated measurement of precisely the gap this page exists to track was missed for 63 days, and was caught now only because an r/LocalLLaMA thread linked the coverage. stateofopensource.ai answers EGRESS_BLOCKED and is new to this repo's blocked list; no first-party read, two search passes (source).

v1.1's figures, all from the one pass that read v1.1 coverage unless noted:

Measurev1.1 value
Open-to-closed capability gap, Mozilla's fit to METR task-horizon data~4.4 months
The same gap, Epoch AI's estimate, which Mozilla says it is in line withfour months
The same gap, Epoch Capabilities Index, from the report's own text (2nd pass)~8 points ≈ four months
Capability doubling time, open3.9 months
Capability doubling time, closed5.5 months
OpenRouter top-10 models by August 2026 token volume that are open-weight8 of 10, 7 of them Chinese-built
Best open model vs closed leader, Artificial Analysis Intelligence Index3 points behind at 60% of the price
Best open model vs Claude Fable 5, same index2 points behind at 30% of the price
The doubling-time pair is the strongest claim and the weakest sourced. If open
capability doubles every 3.9 months against closed's 5.5, the gap is not
stable at four months — it is closing, and the four-month figure is a snapshot of a
converging series rather than a steady state. **One pass, no methodology read, no
second source.** This page does not build on it.

Three cautions, because the numbers are easy to merge and are not the same number. (1) The Artificial Analysis Intelligence Index and the Epoch Capabilities Index are different instruments; "3 points" and "8 points" are not comparable and are not combined here. (2) The OpenRouter count is token volume on one marketplace — a measure of what developers route, not of deployed capability, and nothing read claims otherwise. (3) v1.0's headline figures are two months stale and read almost identically: a gap "narrowed to 3%", an average gap of 3.3 percentage points, costs down up to 50× in three years, a survey of 950+ developers. Those are July's. They are recorded in the snapshot so that a later run does not publish them as September's.

Why it matters here. This page has spent two months arguing about what is being withheld — a licence condition, a capability, an access list, a safety gate. Mozilla is measuring whether the withholding is still load-bearing, and its answer is that on the demand side it largely is not: the models developers actually route to are, by volume, mostly open and mostly Chinese. That is a claim about traffic, not about the frontier, and the two have been diverging all year.

A 7B model argues the gap closes through tool use rather than scale, and publishes the whole stack (2026-09-11, captured 2026-09-17)

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search — ZGCM-1, a 7B dense model trained from scratch on the premise that compact models cannot passively memorise the open web but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use, across a 256K context. It is stated to remain "competitive with frontier models orders of magnitude larger", naming Qwen3-235B-A22B and GLM-5.1, and claims ~4.2× better 16K pre-training time-to-loss (source).

The release list is the longest this page holds: weights from the pre-training, mid-training and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs. Against the two comparable full-stack releases already here — K2 Horizon (2026-09-03, Apache 2.0 on weights and code, with training data, configs and logs) and Hy4 preview (day-one Apache 2.0) — ZGCM-1 adds the data recipes and the training logs, and is smaller than either by two orders of magnitude.

And it does not name a licence. This page's recurring August finding was that the restriction moved into the licence; a release described as "fully open" with no licence stated in anything read is that finding's exact shape. No author, affiliation or lab is named either — for a release whose entire argument is openness, the institution is the one field that cannot be checked without. "Competitive" carries the capability claim and no suite, split, harness or score is attached to it in anything read.

A head of state proposed an open-source bloc, and it is a speech (2026-09-13)

At the BRICS summit in New Delhi on 2026-09-13, Xi Jinping announced that China will lead the creation of an "open-source zone for artificial intelligence" for BRICS countries, as part of an "initiative on open-source and inclusive AI", with stated commitments to cooperation on developing and applying large language models and specialised research and training on AI for BRICS countries. The stated objective is an "open ecosystem" in which members cooperate rather than depend on technology controlled by a handful of companies and countries (2 passes) (source).

No text, charter, budget, timetable, membership list, governing body or legal instrument appears in anything read. It is recorded as a speech commitment, which is what it is — and it is recorded here rather than only on AI Governance because this page has documented, since July, that the open-weights split became formal and geopolitical. Every instrument on that track so far has run through export controls, sanctions and licence terms. This is the first proposal to organise the open side as a bloc, and it comes from the state whose labs supply seven of the eight open models in Mozilla's OpenRouter count above. The adjacency is this wiki's; nothing read connects the two.

Weights were never the whole thing being withheld, and one release just published the rest (2026-09-03)

Every entry below this one argues about weights: whether they ship, when, and on what licence. K2 Horizon moves the axis. IFM (Institute of Foundation Models (IFM), at MBZUAI) released six models, 0.9B to 375B, under Apache 2.0 covering weights and code — and published alongside them the training data (where redistribution licences permit), intermediate checkpoints, training configurations, fine-grained logs and evaluation results. For restricted datasets it publishes source descriptions, construction methods and mixture recipes instead (source).

This is strictly more than any release this page holds. The August argument ran whether → when → terms, and the most permissive answer it reached was Hy4 preview's Apache 2.0 weights. Reproducing a model needs the recipe as well, and until now no release on this page shipped one.

ReleaseWeightsCodeDataRecipe / logs
K2 Horizon (2026-09-03)yes, Apache 2.0yeswhere licensableyes
Hy4 preview (2026-08-28)yes, Apache 2.0not readnono
GLM-5.3 (2026-08-28)yes, bespoke licence + revenue conditionnonono
Muse Glimmer (2026-08-10)yes, Apache 2.0nonopartial (distillation source named)
What it does not publish is the scorecard, and that inverts the usual
omission. Every other release above ships benchmark numbers and withholds the
recipe; this one ships the recipe and withholds the numbers. IFM claims "top-tier performance in every size class", state of the art at 0.9B, 3.7B and 7B, and
matching or beating open-weight MoE models up to 2.6× its size — with **no
benchmark name and no score** attached to any of it in anything read. A release
optimised for reproducibility that gives nobody a number to reproduce is an odd
shape, and this page records it rather than assuming the announcement had a table
the two search passes missed.

Whether it moves the ecosystem is a separate question and this page's own data says probably not much. The 83% / 1% download split holds that models under 1B take 83% of all-time downloads while everything above 100B takes 1%. K2 Horizon's 0.9B, 3.7B and 7B sit in the fat end of that distribution; its 375B flagship sits in the 1%.

The distribution layer changed hands, and this time it is a filing (2026-09-02/03)

The entry below records the acquisition at reporting confidence, with two outlets contradicting each other. It is confirmed: NVIDIA entered a definitive agreement on 2026-09-02 to acquire Hugging Face for ≈ $12.93 billion, disclosed in a Form 8-K, expected to close in the first half of 2027 subject to regulatory approval (source).

Why it is on this page and not only on the two entity pages: the 83% / 1% figure this page reasons from is data the Hub publishes about itself, and the Hub is being bought by a company that sells the compute open weights run on. Jensen Huang states the platform will remain open to the entire AI ecosystem and that NVIDIA compute is not required to use it. Those are commitments in words; nothing read attaches them to a term of the agreement or to any mechanism that outlasts the person who made them.

The largest open release yet carries the fewest conditions, and it is Apache 2.0 (2026-08-28)

Tencent released Hy4 preview — 770B total parameters, 49B active — and open-sourced it the same day, under Apache 2.0, covering both the BF16 checkpoint and an FP8 quantisation, mirrored to ModelScope, GitCode and CNB (source).

Read it against the fortnight it landed in, because this page has spent that fortnight recording conditions. Three Chinese open-weight releases inside fifteen days, and the licences separate more cleanly than the models do:

ModelWeightsLicenceCondition attached
GLM-5.3-Flash2026-08-26, day oneMITnone
Hy4 preview2026-08-28, day oneApache-2.0none
Qwen3.8-Flash-Next2026-08-26, on a countdownqwen-community-1.0licence terms not read here
GLM-5.32026-08-28, after a two-week gateglm-5.3security review for operators above $10B group revenue
The largest model in the group carries the loosest terms. Every argument this
page has recorded for attaching a condition — the cyber-capability disclosure
Z.ai volunteered, the territorial exclusions
MiniMax H3 shipped with, the safety gate that held GLM-5.3 back two
weeks — scales with capability, and Hy4 preview is the biggest artefact of the
four by a factor of four in total parameters. Tencent published it flat. **That
is a data point about what labs choose, not evidence that the conditions are
unnecessary**: no third party has measured this model at all, so nobody outside
Tencent yet knows what was released.

The architecture is the other half of the open-weights argument working as intended. Hy4 preview's attention module is named Gated DeepSeek Sparse Attention — a Tencent frontier model built on a mechanism DeepSeek published. This page has repeatedly recorded open weights as a distribution question; here the reuse is at the level of the design.

A safety gate closes on schedule, and the licence turns out to be the second condition (2026-08-28)

GLM-5.3's weights were published on 2026-08-28, closing the two-week window Z.ai attached to the 2026-08-14 launch — a window this page recorded as the first case it had seen of a Chinese lab delaying open weights behind a stated safety evaluation rather than shipping them with the API. The window closed on the date named. zai-org/GLM-5.3 carries 141 FP8 .safetensors shards and a BF16 sibling at 282 (source).

Two things this page had recorded as unknown are now answered from a first-party artefact, and both answers are more interesting than the release.

What the gate was for. The model card volunteers it: "As we scaled post-training, cyber capability developed faster than we expected", with the model "state of the art on CyberGym for vulnerability discovery" and gains "largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks". That is a lab publishing capability growth it did not plan for as the reason for a release delay — the shape a responsible-scaling disclosure takes, from a lab outside the frameworks Preparedness Framework tracks. What the evaluation consisted of, who ran it and what it concluded are not stated.

The licence is not MIT, and its sibling's was. GLM-5.3 ships under glm-5.3: an MIT-shaped grant plus one clause requiring any "Model as a Service" operator whose group revenue exceeds $10 billion over any consecutive 12 months to pass Z.AI's security review before commercial use, scope "reasonably determined by Z.AI". Model as a Service is defined as third-party inference or fine-tuning access with "meaningful control over the inputs, parameters, or training data", and excludes end-user products embedding the model in features or harnesses, and mere request relaying. GLM-5.3-Flash shipped MIT two days earlier.

This is the fourth distinct shape "open weights" has taken on this page, and the first that gates on who you are rather than on what the model is or when it ships. A revenue threshold at $10B selects hyperscalers and the largest model vendors and nobody else — a self-hoster, a startup or a lab is untouched by clause 2 on its face. Read against Llama's own historical MAU threshold it is a familiar instrument pointed in an unfamiliar direction: the party reserving approval rights over the largest Western serving platforms is a Chinese lab.

And it resolves a tension this page recorded in the direction that empties the gate. GLM-5.3 also carries Z.ai's "Cybersecurity Trusted Access" tier, under which "the model's most sensitive offensive capabilities are reserved exclusively for verified users". Nothing in the card, the licence or the repository mentions that programme. A verified-access control on offensive capability does not survive publication of the weights it gates — anyone may now run the model without passing through Z.ai at all. The question is sharper than when it was first recorded and no less open.

The distribution layer is reported to be changing hands (2026-08-26/27)

Every argument on this page — restriction, licensing, who publishes below 70B, the 83% / 1% download split — is an argument about models that are published somewhere, and in practice that somewhere is Hugging Face. On 2026-08-26/27 two outlets reported NVIDIA acquiring it: The Information that a deal is agreed at $12.9 billion, Business Insider that the two are in talks at more than $13 billion with no deal reached and the possibility it falls apart. Neither company has confirmed anything, and the contradiction is recorded on Hugging Face rather than resolved here (source).

Why it belongs on this page rather than only on a company page. This page's single strongest piece of evidence is the download distribution — models under 1B taking 83% of all-time downloads while everything above 100B takes 1% — and that figure is published by the Hub about itself (source). It was already a self-reported measure. Under new ownership by a company whose business is the hardware those models run on, it would be a self-reported measure with a commercial interest attached. That is a caveat on the evidence, not a prediction about anyone's conduct, and it is the only conclusion the reporting supports.

Nothing read addresses hosting terms, licensing, or the Hub's neutrality between model providers. Those are the questions that would actually change this page, and no source answers them.

An open release used as architecture disclosure, and an open checkpoint used as everyone's measuring stick (2026-08-26)

Two items from one day, and they are different uses of the same word.

Alibaba staged Qwen3.8-Flash-Next for open release with an explicit architectural purpose. The ModelScope teaser describes it as an open-weight multimodal MoE built on "the next-generation architecture that will power the upcoming Qwen4 family", and explicitly not Qwen4 itself — early access to the architecture ahead of the flagship, in standard and FP8, timed for 2026-08-26 23:00 UTC+08:00 (source).

Every open release this page records has been a product published openly. This is a release whose stated function is to let the ecosystem build tooling against an architecture before the model that matters ships — which is a use of open weights as a coordination mechanism rather than as a distribution one. The licence is unknown, as it has been for Qwen 3.8 Max since its weights shipped 2026-08-12, and no benchmark figure exists for it at all. The parameter figures in circulation disagree and are recorded as a conflict on the model page, not as fact.

Meanwhile GPT-OSS 120B did two unrelated jobs in one day. It is one of the three open models OpenAI benchmarked its Jalapeño inference ASIC against (source), alongside DeepSeek R1 670B and Moonshot's Kimi K2.5 — and it is the starting checkpoint for Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs (arXiv:2608.20953), whose GPT-OSS 120B → 60B → MXFP4 pipeline produced the open-weight Hypernova-60B, reported to match or beat its own bfloat16 source on 7 of 9 benchmarks at ~4× less weight memory (source).

That is the argument for open weights that no policy essay on this page makes: a chip vendor and a compression lab could not have compared anything without a checkpoint they both had. The shared measuring stick is the externality, and it is invisible in every adoption number this page counts.

The ecosystem got counted, and the thing being fought over is 1% of it (2026-08-16, captured 2026-08-24)

Every entry below argues about frontier open weights — who may publish a trillion-parameter model, under what licence, with which capability gated. Hugging Face's State of Open Models: Summer 2026 puts a denominator under that argument for the first time, from the platform where the downloads actually happen (source).

Among models that declare a parameter count, those under 1B take 83% of all-time downloads and everything above 100B takes 1%. Public model repositories went 2.43 million → 2.96 million between January and August 2026, datasets crossed 1 million, Spaces 1.00 million → 1.44 million. The post's own illustration of the split is that trillion-parameter Chinese models take the headlines while a small sentence-embedding model is pulled nearly 1.6 billion times.

Read against this page, that is uncomfortable in both directions. The security case for restricting open weights is made about the 1% — the models with frontier agentic and cyber capability — so a policy aimed there touches almost none of the actual usage, which is the strongest available argument that the restriction is narrow rather than an attack on open source. It is equally the strongest argument that the restriction fight is not where the ecosystem's value is, and that both camps have been arguing about the tail. Neither reading is Hugging Face's; the post reports the split without drawing a policy conclusion, and so does this page.

Local inference has a clear leader and it is Chinese. GGUF downloads per month: Qwen 39.6 million, Gemma 20.8 million, Llama 7.5 million — Qwen at ~1.9× Gemma and ~5.3× Llama, with 151,000+ derivatives built on Qwen models, more than any other family. → Alibaba / Qwen AI Lab

And publishing strategy splits by lab in a way this page has been recording release by release. Moonshot, MiniMax, Xiaomi and Z.ai publish almost nothing below 70B, while Tencent and Alibaba Qwen cover the whole range from under 1B upward. That is the difference between opening weights as a capability claim and opening them as a distribution strategy, and it explains why Moonshot AI, MiniMax and Z.ai appear on this page only in frontier-release entries: they publish nothing else.

On licences the post is consistent with what this page holds: Gemma 4, gpt-oss-120b, GLM-5 and most Qwen variants are Apache 2.0 or MIT with no meaningful restrictions, while Llama 4 carries a community licence triggering commercial terms above 700 million monthly active users.

Capture note, because it is a fact about this pipeline rather than the field. This post was reached for and missed on nine consecutive runs (2026-08-15 to 2026-08-23) and carried to two weekly lints; huggingface.co is blocked at CONNECT from the run's sandbox and search had not surfaced enough body to cite. What it does not answer: how many models omit a parameter count and whether excluding them biases the 83%/1% split, and whether GGUF pulls are deduplicated across quantisations — a family shipping more quants would accumulate more downloads for the same adoption.

Open weights got runnable on one machine and the memory to run them got 5× dearer, in the same week (2026-08-20)

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157) reports serving 20+ MoE models across hardware from an 8GB laptop GPU to a single workstation GPU — 35B on a laptop, 284B on a gaming desktop, and GLM-5.2 at 753B on a single workstation GPU — by refusing a fixed offloading strategy and remapping computation and model state onto whatever resources are free, on workloads that include real coding and tool-using agents (source).

That is the strongest capacity datapoint this page holds: a model served commercially by dozens of providers, running on one machine. Every figure in it is a capacity claim and none is a performance claim — no throughput, no latency, no quantisation level, no accuracy check against a datacenter deployment of the same weights. Fitting a model is not serving it.

The economics moved the other way the same week. Consumer DDR5 is up as much as 485% year over year; a 128GB kit is $3,399, roughly 10× its lowest tracked price of about $329, with mainstream 64GB kits past $1,000. The cause given is AI infrastructure demand, with hyperscale customers reported to be reserving much of the industry's future DRAM capacity, and no forecast read expects relief before late 2027 (source).

FreeToken's method is precisely to substitute host memory and bandwidth for GPU capacity. So the technique that makes frontier open weights runnable at home arrived in the quarter when the resource it spends became the scarce one. Neither source mentions the other; the pairing is this wiki's. The figure that would settle how much it matters — host RAM required per hardware tier — is not published in either.

One vendor, two licences, eleven days apart (2026-08-17)

MiniMax Music 3.0's weights were published on 2026-08-13 under the MiniMax-Music3 Community License, reported to carry no territorial exclusion and to permit downloading and running the model rather than only calling a hosted endpoint (source).

Eleven days earlier the same company shipped MiniMax H3 under the H3 Community License, which excludes the United States, European Union, United Kingdom and South Korea from its Applicable Territory, citing generative-video regulation in those jurisdictions (source).

Same vendor, same month, same phrase on the announcement, and a reader in London may use one model and not the other. This page has recorded the open/closed axis (weights or no weights) and the staging axis (GLM-5.3's weights withheld pending evaluation). This is a third: jurisdiction, decided per release rather than per company. The direction is also worth noting — the permissive licence is the later one, so the H3 carve-out reads as a response to video regulation specifically, not a corporate policy tightening.

Permissible training data, benchmarked (2026-08-18)

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517) is the first artefact on this page to argue that the input side of openness is affordable. A 1B HRM-architecture model trained from scratch on 161 datasets of permissible data is reported to outperform HRM-Text 1B, to compete with Qwen 3.5 4B and Gemma 4 E2B across 20 benchmarks, and to set a new state of the art for Danish (source).

The constraint in its title covers post-training data. Whether the pre-training corpus is under the same restriction is not stated in anything read, and the narrow reading — that the post-training stage can be made permissible cheaply — is the one this wiki publishes. Everything else on this page concerns what a release permits; this is the only entry concerning what went in.

Two Chinese labs shipped on the same day and took opposite halves of the deal (2026-08-14)

The clearest natural experiment this page has carried. Same day, same country, same claim to open weights, opposite artefacts.

Alibaba / Qwen AI Lab shipped the weights and named the licence. Qwen 3.8 27B is out — 27.78B dense, Apache 2.0, on Hugging Face and ModelScope, with an FP8 build alongside (source). This closes the eleven-day gap between an announcement that named no licence and an artefact that has one. Four days late, and Alibaba never restated, moved or withdrew the 2026-08-10 date it missed.

Z.ai shipped the claim and withheld the weights. GLM-5.3 launched to GLM Coding Plan subscribers from $18/month, marketed as "the strongest open-weights coding model", with open weights and API access stated for roughly two weeks out, in stages, after a safety evaluation (source). Nothing is downloadable. For comparison, GLM-5.2 put subscriber access and open weights three days apart in June — so this is a change of practice by the same lab, not a lab that always worked this way.

Three things follow.

  1. "Open-weights" has become a description of a roadmap, not of a file. The GLM-5.3 headline uses the term for a model whose weights do not exist publicly. That is a usage this page should track, because it is the point at which the label stops distinguishing anything.
  2. A safety evaluation is now given as the reason for the gap. This is the first Chinese release captured here to condition a weights drop on one. Read generously it is the staged-release norm arriving; read plainly it is also a two-week paid-subscription window with the open label attached. Both readings fit what was published and nothing read separates them.
  3. The announce-then-open window is measurable again. Qwen3.8-27B ran 2026-08-03 → 2026-08-14, eleven days, against the 24 days measured for Qwen 3.8 Max below. GLM-5.3's window is stated but not yet elapsed — the first entry on this page that is a promise rather than a measurement, and worth returning to around 2026-08-28.

The opened checkpoint measures the same as the API it was cut from (2026-08-16)

Artificial Analysis's 2026-08-16 leaderboard carries both forms of Qwen 3.8 Max as separate rows, and they land on the same number (source):

ModelIntelligence IndexContext WindowCost per Task USD
Qwen3.8 Max581M$1.13
Qwen3.8 2.4T A95B58984k$1.09
Neither row exists in the 2026-08-09 snapshot for the open checkpoint, so this is
the first independent side-by-side this repo holds.

Why it is worth recording even though it is the expected result. This page has spent three weeks on the terms of opening — the 24-day window, the licence, the promise. What none of that established is whether the artefact a lab hands over is the artefact it was selling. Here it is: same index, at $1.09 against $1.13. "Open weights" for this release is not a capability discount, and the paid window bought time rather than quality.

The one difference is instructive and is not a capability gap: 984k against 1M context. A hosted flagship states its context window; a self-hosted checkpoint is measured at whatever the serving deployment allocated. What a reader gets from an open checkpoint is set by whoever is running it — which is the same "open ≠ runnable" point the section below makes about scale, arriving here as a measurement artefact rather than a hardware bill.

Held against GLM-5.3, whose weights are still stated as roughly two weeks out behind a safety evaluation, the contrast the 08-14 entry drew is now quantified on one side and still unquantifiable on the other.

A Max-class model opens, and "open" separates from "runnable" (2026-08-12)

Alibaba / Qwen AI Lab published open weights for Qwen3.8-2.4T-A95B — the open-weight form of Qwen 3.8 Max — and for Qwen 3.8 27B, the first time a Max-class Qwen has been opened (source).

Two things on this page change because of it.

First, the announce-then-open pattern completed and can now be timed. Coverage of the release names the sequence directly: Chinese labs "have spent 2026 converging on the same play: launch closed with a benchmark case, monetize the API window, then open the weights once the news cycle has done its work" (source). For this model the window was 24 days — closed preview 2026-07-19, paid API 2026-08-03, open weights 2026-08-12. That is a measurement rather than a characterisation, and it is the first one this page carries. MiniMax H3 ran the same sequence and Kimi K3 ran a shorter version of it.

Second, scale has decoupled "open" from "usable", and it did so at both ends. The released flagship is 4.9 TB at full precision; the smallest quantisation anyone has built is 397 GB, an 89–91% reduction that still does not reach a single consumer GPU (source). Alibaba shipped the 27B alongside it precisely because of that gap. So an open release at the frontier now means inspectable and fine-tunable by well-resourced parties, not self-hostable by individuals — which cuts across both sides of the argument this page tracks. The security case for open weights (many eyes, containment) survives it; the democratisation case largely does not.

Third, and separately: the licence has not been named. Weights are downloadable and no source read states the terms (source). Earlier Qwen lines shipped Apache-2.0, which is precedent and not a licence. This page has already recorded MiniMax H3 shipping under terms that excluded the four jurisdictions writing generative-AI rules; an unnamed licence on the largest open release of the year is the same question left open rather than answered, and it is the question the policy fight actually turns on.

Answered for one of the two, on 2026-08-14. Qwen 3.8 27B shipped under Apache 2.0 (source) — so the precedent held, which is worth recording precisely because this page declined to assume it. The flagship's licence is still unnamed in anything read. Note also that this paragraph's original framing was too generous to the 08-12 drop: the 27B did not ship that day, as the correction on its model page records. Only the Max-class weights did.

Meta ships the student open and keeps the teacher closed (2026-08-10)

Meta AI released Muse Glimmer under Apache 2.0 with open weights — a 30B dense multimodal model built to run offline on a single consumer GPU — five days after Muse Spark 1.2 shipped as a closed paid API on the same product line. Glimmer is distilled from Muse Spark: pre-training used logit distillation on the teacher's outputs (source).

The structure is worth naming, because "Meta returns to open source" is how the release was covered and it is not quite what happened. The frontier model stays closed and the derived small model goes open. That is a third position alongside the two this page has been tracking — not the Llama-era open frontier, and not the closed turn of spring 2026 — and it costs the vendor nothing at the frontier while supplying the policy argument in full.

Mark Zuckerberg supplied that argument explicitly on release day: US policy must reduce friction if American open source models are to lead, with DeepSeek and Moonshot named as getting "uncomfortably close to the American frontier" (source). The open/closed dispute this page records as institutional is here being argued as national, which is the same move Anthropic's testing-mandate position and China's distillation counter-claim made from their own directions.

The alliance ships its first member model, and it is a guardrail (2026-08-04)

Mistral AI released Shieldstral 1.0 on 2026-08-04 as an inaugural Open Secure AI Alliance member release, alongside NVIDIA — a 3B policy-adaptive multimodal safety classifier under Apache 2.0, running on a single 16GB GPU, reported at 84.9% average F1 on text safety and 83.8% on multimodal safety against OmniGuard-7B's 77.6% and LlavaGuard-7B's 71.6% (source) (Mistral).

It does not answer the open problem below, and it is worth being exact about why. The alliance's unmet claim is that defenders need frontier-class open models; a 3B classifier is not one, and nothing about this release moves Inkling's 41 on the Artificial Analysis Intelligence Index closer to Claude Opus 5's 61.

What it does is answer a different part of Huang's case. His stated complaint was that closed models blocked forensics during the Hugging Face incident — that the defensive layer was unavailable when it was needed. A guardrail model published under Apache 2.0, runnable on hardware a small operator already owns, is that layer being shipped rather than argued for. It is the first alliance output here that a defender can actually deploy; NOOA, the launch contribution, is a framework for testing agents, not a model.

The design is the part that bears on this page. Shieldstral's harm taxonomy is supplied at inference time as natural-language policy rather than trained into the weights, so a deployer re-aims it at its own rules without retraining (source). Every argument on this page about who decides what a model refuses has assumed that decision lives in the weights and therefore with whoever trained them. Here it does not. Whether that generalises past classification to generation is untested and nothing read claims it does — but it is the first published counter-example to an assumption the whole dispute rests on.

Open Secure AI Alliance launched (2026-07-27)

NVIDIA announced the Open Secure AI Alliance on July 27, 2026 — an industry body to build and share open models and tools for AI defenders (source) (NVIDIA).

  • Governance: under the Linux Foundation umbrella, building on the Foundation's Akrites vulnerability-disclosure effort and existing OpenSSF work (Linux Foundation)
  • First technical contribution: NOOA (NVIDIA-labs OO Agents), an Apache 2.0 research framework for testing, tracing, auditing and governing agent behavior, on GitHub at launch
  • Scope: the full agent stack — identity, permissions, isolation, guardrails, logs, model formats, multi-model scanning, secure coding workflows
  • Founding partners include: Microsoft, IBM, Red Hat, Hugging Face, Mistral, Cloudflare, CrowdStrike, Palantir, Databricks, GitHub, LangChain, Perplexity, Nous Research, Reflection AI, Thinking Machines Lab, SpacexAI, vLLM, SAP, Siemens, SK Telecom, NAVER, the Linux Foundation
  • Absent: OpenAI, Anthropic, Google, Meta and Amazon — between them the builders of most frontier models the alliance says defenders need (TNW)

Jensen Huang's stated case, quoted in the announcement: "Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community." And, pointedly: "During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion."

The forensics claim is documented, not rhetorical

Huang's second sentence has a primary source behind it. Hugging Face's technical timeline of the July 2026 intrusion, published July 27, records that the response team tried to analyze attack artifacts with commercially hosted LLMs and found the models' safety guardrails blocked analysis of prompts containing genuine attack artifacts — forcing them onto a self-hosted, open-weight model (source) (HF).

This is the strongest empirical argument the open-weights side has produced: not that open models are safer, but that closed models are unavailable precisely when defenders need them. → AI-Enabled Cyberattacks

Anthropic states its position (2026-07-28)

After a week in which community reporting held that Anthropic was lobbying for open-weight restrictions, Dario Amodei responded on July 28 (source) (Bloomberg):

  • Anthropic has never advocated a ban on open-weight models
  • Open-weight models without dangerous capabilities are a public good
  • Three measures proposed instead: (1) tighter export controls on advanced AI chips and chipmaking equipment to China; (2) a crackdown on industrial-scale model distillation; (3) mandatory safety testing for sufficiently capable models, open or closed

The third measure is where the disagreement actually lives. A capability-triggered testing mandate is not a ban, but it applies to a release the moment weights leave the building — and the open-weights camp's objection is that no open project can satisfy a pre-release testing regime the way a lab with a safety org can. → Anthropic

China reframes distillation as a two-way practice (2026-07-28)

China's Ministry of Commerce answered US sanctions threats the same day, stating that "many American artificial intelligence enterprises have distilled Chinese models during their research, development and training processes", calling the US accusations groundless and "AI hegemonism", and warning of countermeasures (source) (The Register).

This answers Treasury Secretary Scott Bessent's July 21 threat of sanctions and Entity List blacklisting over industrial-scale distillation. With Chinese labs now shipping the largest open-weight models (Kimi K3, DeepSeek V4, Qwen 3.8 Max), "restrict open weights" and "restrict Chinese models" have become close to the same policy. → AI Governance, Moonshot AI, DeepSeek

Two Western open-weight releases the policy fight did not cite (2026-07-15, 2026-07-21)

While the argument above ran, two Western labs shipped open weights. One of them, Thinking Machines Lab, is a founding partner of the alliance — its weights landed twelve days before the launch and neither the announcement nor the coverage named them.

This bears directly on the Open Secure AI Alliance problem below — that the alliance proposes frontier-class open models for defenders and none of its members ships one. The Inkling release is the nearest any member has come, and it does not close the gap. Neither model is frontier-class on the one third-party measure this wiki holds: Inkling scores 41 on the Artificial Analysis Intelligence Index read 2026-08-02 against Claude Opus 5 (max) at 61, and Laguna S 2.1 appears in none of that table's 260 rows (source).

What they do change is the geography of the argument. "Restrict open weights" and "restrict Chinese models" had become close to the same policy because the largest open-weight releases were Chinese; two US-addressed releases inside a week, one under Apache 2.0, put a Western constituency on the side that a distillation-focused enforcement regime would burden. Coverage framed Laguna S 2.1 in exactly those terms — "the West's answer to DeepSeek and Qwen" (source).

The restriction moved into the licence (2026-08-03)

Two Chinese releases on the same day showed the dispute above being settled privately, by licence terms and checkpoint selection, rather than by any of the policy mechanisms the alliance and the labs are arguing over.

MiniMax H3 — weights published, four jurisdictions excluded. MiniMax shipped H3-Base to Hugging Face on 2026-08-03 under the MiniMax Community License, which is free below $20M revenue with UI attribution and excludes the United States, European Union, United Kingdom and South Korea by territory. MiniMax's stated reason is that those regions "are currently developing or enforcing AI-related regulations that may have specific implications for generative video models"; organisations there may apply for a formal licence (source).

This is a shape neither camp's argument anticipated. The policy fight assumes a government deciding whether weights may be released; here the releasing lab excluded the regulating jurisdictions itself, pre-emptively, and the excluded set is precisely the regions writing generative-AI rules. Weights that are downloadable worldwide and licensed nowhere in the West are neither the "open weights for defenders" the alliance argues for nor the gated release the labs propose.

And the released artefact is not the marketed one. H3-Context-IR and H3-Regenerate-2K were withheld, so the 1440p output the model launched on cannot be reproduced locally — MiniMax's own full workflow calls its hosted API for the two missing modules (source). "Open weights" named a 768p subset of the product. Compare Inkling (Apache 2.0) and Laguna S 2.1 (OpenMDW-1.1), which restrict neither by territory nor by component.

Qwen 3.8 Max — the date finally attached. Alibaba made Qwen3.8-Max generally available on 2026-08-03 via API and stated open weights for the week of 2026-08-10, alongside Qwen 3.8 27B (source). The "open-weight soon" claim of 2026-07-19 had gone 15 days without one. Both of the largest announced Chinese open releases of 2026 have now shipped an API first and the weights later or not yet — a sequence worth recording, because the policy argument treats "open-weight lab" as a standing property rather than a promise with a date.

Measurement arrived the same day. Interconnects launched an Artifacts Hub and an Adoption Dashboard tracking Hugging Face downloads and derivative-model counts by geography and organisation, updated daily and broken out USA vs China vs EU, using Relative Adoption Metrics that normalise downloads within a size class (source). The US–China gap is the framing, and it is the quantity this section has repeatedly had to assert from release counts rather than measure.

2026-08-16 — Amodei's position moves from "not a ban" to a claim about where power actually sits. This page has carried Amodei's 2026-07-28 statement as the reference point: Anthropic "never advocated" a ban, open-weight models without dangerous capabilities are a public good, and the ask is export controls, anti-distillation enforcement and mandatory testing regardless of open or closed (source). A long post on X three weeks later adds an argument that statement did not make (source):

2026-07-282026-08-16
Open weights are not to be bannedOpen weights are "somewhat better" on power concentration — and shift it to whoever holds the most compute and chips
Three measures: export controls, distillation, testingA FINRA-like entity — a self-regulatory organisation
Framed as a rebuttal to a ban accusationFramed as a false choice between regulatory capture and wide distribution
The load-bearing move is the second row of the first column: it answers the
strongest argument for open weights — that distribution is itself the check on
concentrated power — not by denying it but by relocating it. If the binding
constraint is compute rather than weights, then opening a model redistributes
something that was not scarce.

Two things keep this from being adopted as settled. It is made by the CEO of a lab that ships no open weights, which is the most interested position from which to make it. And the compute-concentration claim is stated, not measured — nothing read attaches a number to how much of the derivative ecosystem is gated on capacity rather than on access to weights. The Adoption Dashboard above is the first instrument on this page that could begin to answer it, and nobody has pointed it at this question.

The post itself was not read. x.com is blocked from this environment and no source read gives its status URL, so the whole entry rests on third-party reporting — one rank below where a lab CEO's own post would sit under this repo's trust_order (source).

Open Problems

  • Does open weights still transfer the capability? (added 2026-08-21) Z.ai's CEO is reported to argue that memorization prefers parameters while reasoning and advanced skills come from post-training — with GLM-5.3 as the demonstration: GLM-5.2's base untouched, MIT-licensed and downloadable since 2026-06-16, every gain from post-training that is not being published (source). This page has treated the licence and the weight-drop date as the axis of the dispute. If the claim holds, a lab can publish the base, satisfy every open-weights commitment tracked here, and retain the half that produced the capability — and GLM-5.3's staged release is shaped exactly that way. Nobody has said whether that is the intent, and no independent measurement of the claim exists. → Post-Training Scaling
  • Provenance is now checkable without the publisher's cooperation, and it does not check the thing this page argues about (added 2026-08-21). Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929) verifies shared weight ancestry from checkpoints alone — data-free, white-box, AUROC = 1.0 separating fine-tuned, LoRA-merged, pruned and quantized descendants from independent models, unchanged under function-preserving laundering, 76× faster than the nearest robust baseline. That makes licence compliance auditable: whether a published checkpoint is secretly a fine-tune of an encumbered base is exactly weight ancestry. But distilled models group with the independent ones, correctly by the method's own definition — it measures weight ancestry, not behavioral similarity. So it settles nothing about the Alibaba / Qwen AI Lab distillation allegation or MOFCOM's counter-claim below, and a reader who saw only "AUROC = 1.0 for lineage" would conclude the opposite.
  • Is compute concentration measurable against weight access? Amodei's 2026-08-16 claim that open weights shift power to those with the most chips is the sharpest version of the argument on this page and the only one with no number attached. The Adoption Dashboard tracks downloads and derivatives; nothing tracks who could actually serve what they downloaded.
  • What exactly triggers a testing mandate? Amodei's proposal turns on "sufficiently capable", the same undefined threshold that has blocked the US voluntary framework's "covered frontier model" designation for months.
  • Who tests a model with no owner? Mandatory pre-release testing assumes a releasing entity with the resources to run it. Fine-tunes and re-releases of open weights have no such party.
  • The alliance has no frontier model. Its stated mission is to give defenders frontier-class open models; none of its members currently ships one at the level of the labs that declined to join — Mistral and the Chinese labs are the nearest, and the latter are the ones under sanctions threat. Founding partner Thinking Machines Lab shipped Apache 2.0 weights twelve days before the launch and it did not change this: Inkling scores 41 on the Artificial Analysis Intelligence Index against Claude Opus 5 (max) at 61 (source). Mistral's Shieldstral 1.0 (2026-08-04) does not change it either — a 3B classifier is not the frontier model in question — though it is the first member release a defender can deploy.
  • Is an open guardrail worth more than an open frontier model to a defender? Shieldstral makes the question answerable rather than rhetorical: the alliance's stated need is frontier-class weights, but its first shipped model is a small classifier, and nothing read argues which does more for the defenders it names. Watch whether the mission statement moves toward what members actually ship.
  • Symmetry cuts both ways. If distillation is as universal as MOFCOM claims, an enforcement regime against it constrains US labs' training pipelines as much as Chinese ones — which no US proposal has yet acknowledged.
  • Is "closed models refuse forensics" fixable without opening weights? A carve-out for verified incident responders would answer Huang's specific complaint without conceding the general argument. No lab has proposed one.

Key Papers

  • Frontier Pacing — the countervailing thread: a pacing mechanism presumes a frontier held by a countable number of coordinatable actors

  • Military and Intelligence Capability Evals — 2026-09-10: the first measurement on this wiki of what open weights actually hand over in targeting and weapons tasks, against which "behind the frontier" becomes checkable

  • AI Governance — the regulatory frame this dispute sits inside

  • AI-Enabled Cyberattacks — the July 2026 intrusion both camps cite

  • Agents (LLM Agents) — agent harnesses are what the alliance proposes to secure

  • NVIDIA — convener of the alliance

  • Anthropic — the position most often characterized as restrictionist

  • OpenAI — absent from the alliance; the lab whose evaluation the intrusion escaped

  • Moonshot AI — largest open-weight release to date, and a named target of distillation sanctions

  • Thinking Machines Lab — Apache 2.0 frontier-scale weights, 2026-07-15

  • Poolside — OpenMDW-1.1 agentic coding weights, 2026-07-21

  • MiniMax — first territory-excluded open-weight licence recorded here, 2026-08-03

  • Alibaba / Qwen AI Lab — API first, weights dated 2026-08-10

  • Mistral AI — first alliance member to ship a model under the alliance banner

Conflicting Reports

Three leaderboard cells this page carries read differently in a second snapshot it also cites, and every one of them is a read-date difference rather than a disagreement. Recorded here because the schema wants a cross-source difference written down rather than reconciled, and because a checker that pairs a row label with a column heading has no notion of when a board was read.

CellThis pageOther cited snapshot
Nemotron 3 Ultra · Artificial Analysis Intelligence Index23, read 2026-09-1338 (2026-08-02)
Qwen3.8 Max · Cost per Task USD$1.13, read 2026-08-16$2.67 (2026-09-13)
Qwen3.8 2.4T A95B · Cost per Task USD$1.09, read 2026-08-16$2.16 (2026-09-13)
The first row is the index revision described in the 2026-09-21 section above, and the other two are prices moving over four weeks. Neither is a contradiction and neither figure was changed: each is correct for the snapshot it cites, and this page states in that section that no figure from one reading may be set against a figure from another.

Founding member count. Outlets published different numbers on the same launch: "30+ companies" (Tom's Hardware), "37-member" (CoinDesk, The Hacker News), "more than 40 founding members" (MLQ News), "44 founding firms" (AI Weekly). NVIDIA's own announcement lists the partners by name without stating a total (NVIDIA); the roster reproduced in the snapshot is longer than any of the reported counts. Unresolved — likely reflects the roster growing between embargo and publication.

Referenced by

2026-W392026-W40Adversarial DistillationAI GovernanceAI-Enabled CyberattacksAI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv:2608.12307)Aleph AlphaAlibaba / Qwen AI LabAn Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsAnt Group (inclusionAI / AntLing)AnthropicAugust 2026 — Monthly DigestClaude Opus 5.5Content Provenance (AI output marking)DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517)Eval Environment ContainmentFeyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber ModelsFreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157)Frontier PacingFugu MaxGLM-5.3Hugging FaceHunyuan-A13B Technical ReportIBMInstitute of Foundation Models (IFM)Iris: Climbing to the Search FrontierJuly 2026 — Monthly DigestJune 2026 — Monthly DigestK2 HorizonKimi K2.8 PreviewKolibri-1Liquid AIMeta AIMilitary and Intelligence Capability EvalsMiniMaxMiniMax H3MiniMax Music 3.0Mistral AIMoonshot AIMuse GlimmerNemotron 3.5 LightningNeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing HarnessNVIDIAOpenAIPoolsidePost-Training ScalingPrismMLQuantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs (arXiv:2608.20953)Qwen 3.8 27BQwen 3.8 MaxQwen-Drive-1.0-4BShieldstral (arXiv:2607.25857)Shieldstral 1.0Stealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867)TencentThinking Machines LabTraining Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929)Unlocking Lossless Speedups in LLMs via Discrete DiffusionVentor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391)Weekly Synthesis — W31 (July 27 – August 2, 2026)Weekly Synthesis — W32 (2026-08-03 → 2026-08-09)Weekly Synthesis — W33 (2026-08-10 → 2026-08-16)Weekly Synthesis — W35 (2026-08-24 → 2026-08-30)XiaomiZ.aiZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Sources