AI Trend Notifier
EN
← wiki

$ cat wiki/concepts/open-weights-policy.md

Open-Weights Policy Fight

Definition

The 2026 policy dispute over whether openly released model weights should be restricted, and on what grounds. The question is no longer academic: it now has an industry coalition on one side, three frontier labs and a Treasury sanctions threat on the other, and a live security incident that both camps cite as evidence.

Why It Matters

  • The industry split is now formal. Until July 2026 the disagreement was rhetorical. The launch of the Open Secure AI Alliance put roughly forty companies behind open weights as a security position, and made the absence of the frontier labs a matter of public record.
  • It is being argued from incident evidence, not preference. Both sides now point at the same July 2026 agent intrusion — one to argue open weights contained it, the other to argue agentic capability needs gating. → AI-Enabled Cyberattacks
  • The distillation dispute is the geopolitical face of the same fight. Restricting open weights and sanctioning Chinese labs for distillation are the same lever pointed in two directions. → AI Governance

State of the Art (2026-08-15)

Open weights got runnable on one machine and the memory to run them got 5× dearer, in the same week (2026-08-20)

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157) reports serving 20+ MoE models across hardware from an 8GB laptop GPU to a single workstation GPU — 35B on a laptop, 284B on a gaming desktop, and GLM-5.2 at 753B on a single workstation GPU — by refusing a fixed offloading strategy and remapping computation and model state onto whatever resources are free, on workloads that include real coding and tool-using agents (source).

That is the strongest capacity datapoint this page holds: a model served commercially by dozens of providers, running on one machine. Every figure in it is a capacity claim and none is a performance claim — no throughput, no latency, no quantisation level, no accuracy check against a datacenter deployment of the same weights. Fitting a model is not serving it.

The economics moved the other way the same week. Consumer DDR5 is up as much as 485% year over year; a 128GB kit is $3,399, roughly 10× its lowest tracked price of about $329, with mainstream 64GB kits past $1,000. The cause given is AI infrastructure demand, with hyperscale customers reported to be reserving much of the industry's future DRAM capacity, and no forecast read expects relief before late 2027 (source).

FreeToken's method is precisely to substitute host memory and bandwidth for GPU capacity. So the technique that makes frontier open weights runnable at home arrived in the quarter when the resource it spends became the scarce one. Neither source mentions the other; the pairing is this wiki's. The figure that would settle how much it matters — host RAM required per hardware tier — is not published in either.

One vendor, two licences, eleven days apart (2026-08-17)

MiniMax Music 3.0's weights were published on 2026-08-13 under the MiniMax-Music3 Community License, reported to carry no territorial exclusion and to permit downloading and running the model rather than only calling a hosted endpoint (source).

Eleven days earlier the same company shipped MiniMax H3 under the H3 Community License, which excludes the United States, European Union, United Kingdom and South Korea from its Applicable Territory, citing generative-video regulation in those jurisdictions (source).

Same vendor, same month, same phrase on the announcement, and a reader in London may use one model and not the other. This page has recorded the open/closed axis (weights or no weights) and the staging axis (GLM-5.3's weights withheld pending evaluation). This is a third: jurisdiction, decided per release rather than per company. The direction is also worth noting — the permissive licence is the later one, so the H3 carve-out reads as a response to video regulation specifically, not a corporate policy tightening.

Permissible training data, benchmarked (2026-08-18)

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517) is the first artefact on this page to argue that the input side of openness is affordable. A 1B HRM-architecture model trained from scratch on 161 datasets of permissible data is reported to outperform HRM-Text 1B, to compete with Qwen 3.5 4B and Gemma 4 E2B across 20 benchmarks, and to set a new state of the art for Danish (source).

The constraint in its title covers post-training data. Whether the pre-training corpus is under the same restriction is not stated in anything read, and the narrow reading — that the post-training stage can be made permissible cheaply — is the one this wiki publishes. Everything else on this page concerns what a release permits; this is the only entry concerning what went in.

Two Chinese labs shipped on the same day and took opposite halves of the deal (2026-08-14)

The clearest natural experiment this page has carried. Same day, same country, same claim to open weights, opposite artefacts.

Alibaba / Qwen AI Lab shipped the weights and named the licence. Qwen 3.8 27B is out — 27.78B dense, Apache 2.0, on Hugging Face and ModelScope, with an FP8 build alongside (source). This closes the eleven-day gap between an announcement that named no licence and an artefact that has one. Four days late, and Alibaba never restated, moved or withdrew the 2026-08-10 date it missed.

Z.ai shipped the claim and withheld the weights. GLM-5.3 launched to GLM Coding Plan subscribers from $18/month, marketed as "the strongest open-weights coding model", with open weights and API access stated for roughly two weeks out, in stages, after a safety evaluation (source). Nothing is downloadable. For comparison, GLM-5.2 put subscriber access and open weights three days apart in June — so this is a change of practice by the same lab, not a lab that always worked this way.

Three things follow.

  1. "Open-weights" has become a description of a roadmap, not of a file. The GLM-5.3 headline uses the term for a model whose weights do not exist publicly. That is a usage this page should track, because it is the point at which the label stops distinguishing anything.
  2. A safety evaluation is now given as the reason for the gap. This is the first Chinese release captured here to condition a weights drop on one. Read generously it is the staged-release norm arriving; read plainly it is also a two-week paid-subscription window with the open label attached. Both readings fit what was published and nothing read separates them.
  3. The announce-then-open window is measurable again. Qwen3.8-27B ran 2026-08-03 → 2026-08-14, eleven days, against the 24 days measured for Qwen 3.8 Max below. GLM-5.3's window is stated but not yet elapsed — the first entry on this page that is a promise rather than a measurement, and worth returning to around 2026-08-28.

The opened checkpoint measures the same as the API it was cut from (2026-08-16)

Artificial Analysis's 2026-08-16 leaderboard carries both forms of Qwen 3.8 Max as separate rows, and they land on the same number (source):

ModelIntelligence IndexContext WindowCost per Task USD
Qwen3.8 Max581M$1.13
Qwen3.8 2.4T A95B58984k$1.09
Neither row exists in the 2026-08-09 snapshot for the open checkpoint, so this is
the first independent side-by-side this repo holds.

Why it is worth recording even though it is the expected result. This page has spent three weeks on the terms of opening — the 24-day window, the licence, the promise. What none of that established is whether the artefact a lab hands over is the artefact it was selling. Here it is: same index, at $1.09 against $1.13. "Open weights" for this release is not a capability discount, and the paid window bought time rather than quality.

The one difference is instructive and is not a capability gap: 984k against 1M context. A hosted flagship states its context window; a self-hosted checkpoint is measured at whatever the serving deployment allocated. What a reader gets from an open checkpoint is set by whoever is running it — which is the same "open ≠ runnable" point the section below makes about scale, arriving here as a measurement artefact rather than a hardware bill.

Held against GLM-5.3, whose weights are still stated as roughly two weeks out behind a safety evaluation, the contrast the 08-14 entry drew is now quantified on one side and still unquantifiable on the other.

A Max-class model opens, and "open" separates from "runnable" (2026-08-12)

Alibaba / Qwen AI Lab published open weights for Qwen3.8-2.4T-A95B — the open-weight form of Qwen 3.8 Max — and for Qwen 3.8 27B, the first time a Max-class Qwen has been opened (source).

Two things on this page change because of it.

First, the announce-then-open pattern completed and can now be timed. Coverage of the release names the sequence directly: Chinese labs "have spent 2026 converging on the same play: launch closed with a benchmark case, monetize the API window, then open the weights once the news cycle has done its work" (source). For this model the window was 24 days — closed preview 2026-07-19, paid API 2026-08-03, open weights 2026-08-12. That is a measurement rather than a characterisation, and it is the first one this page carries. MiniMax H3 ran the same sequence and Kimi K3 ran a shorter version of it.

Second, scale has decoupled "open" from "usable", and it did so at both ends. The released flagship is 4.9 TB at full precision; the smallest quantisation anyone has built is 397 GB, an 89–91% reduction that still does not reach a single consumer GPU (source). Alibaba shipped the 27B alongside it precisely because of that gap. So an open release at the frontier now means inspectable and fine-tunable by well-resourced parties, not self-hostable by individuals — which cuts across both sides of the argument this page tracks. The security case for open weights (many eyes, containment) survives it; the democratisation case largely does not.

Third, and separately: the licence has not been named. Weights are downloadable and no source read states the terms (source). Earlier Qwen lines shipped Apache-2.0, which is precedent and not a licence. This page has already recorded MiniMax H3 shipping under terms that excluded the four jurisdictions writing generative-AI rules; an unnamed licence on the largest open release of the year is the same question left open rather than answered, and it is the question the policy fight actually turns on.

Answered for one of the two, on 2026-08-14. Qwen 3.8 27B shipped under Apache 2.0 (source) — so the precedent held, which is worth recording precisely because this page declined to assume it. The flagship's licence is still unnamed in anything read. Note also that this paragraph's original framing was too generous to the 08-12 drop: the 27B did not ship that day, as the correction on its model page records. Only the Max-class weights did.

Meta ships the student open and keeps the teacher closed (2026-08-10)

Meta AI released Muse Glimmer under Apache 2.0 with open weights — a 30B dense multimodal model built to run offline on a single consumer GPU — five days after Muse Spark 1.2 shipped as a closed paid API on the same product line. Glimmer is distilled from Muse Spark: pre-training used logit distillation on the teacher's outputs (source).

The structure is worth naming, because "Meta returns to open source" is how the release was covered and it is not quite what happened. The frontier model stays closed and the derived small model goes open. That is a third position alongside the two this page has been tracking — not the Llama-era open frontier, and not the closed turn of spring 2026 — and it costs the vendor nothing at the frontier while supplying the policy argument in full.

Mark Zuckerberg supplied that argument explicitly on release day: US policy must reduce friction if American open source models are to lead, with DeepSeek and Moonshot named as getting "uncomfortably close to the American frontier" (source). The open/closed dispute this page records as institutional is here being argued as national, which is the same move Anthropic's testing-mandate position and China's distillation counter-claim made from their own directions.

The alliance ships its first member model, and it is a guardrail (2026-08-04)

Mistral AI released Shieldstral 1.0 on 2026-08-04 as an inaugural Open Secure AI Alliance member release, alongside NVIDIA — a 3B policy-adaptive multimodal safety classifier under Apache 2.0, running on a single 16GB GPU, reported at 84.9% average F1 on text safety and 83.8% on multimodal safety against OmniGuard-7B's 77.6% and LlavaGuard-7B's 71.6% (source) (Mistral).

It does not answer the open problem below, and it is worth being exact about why. The alliance's unmet claim is that defenders need frontier-class open models; a 3B classifier is not one, and nothing about this release moves Inkling's 41 on the Artificial Analysis Intelligence Index closer to Claude Opus 5's 61.

What it does is answer a different part of Huang's case. His stated complaint was that closed models blocked forensics during the Hugging Face incident — that the defensive layer was unavailable when it was needed. A guardrail model published under Apache 2.0, runnable on hardware a small operator already owns, is that layer being shipped rather than argued for. It is the first alliance output here that a defender can actually deploy; NOOA, the launch contribution, is a framework for testing agents, not a model.

The design is the part that bears on this page. Shieldstral's harm taxonomy is supplied at inference time as natural-language policy rather than trained into the weights, so a deployer re-aims it at its own rules without retraining (source). Every argument on this page about who decides what a model refuses has assumed that decision lives in the weights and therefore with whoever trained them. Here it does not. Whether that generalises past classification to generation is untested and nothing read claims it does — but it is the first published counter-example to an assumption the whole dispute rests on.

Open Secure AI Alliance launched (2026-07-27)

NVIDIA announced the Open Secure AI Alliance on July 27, 2026 — an industry body to build and share open models and tools for AI defenders (source) (NVIDIA).

  • Governance: under the Linux Foundation umbrella, building on the Foundation's Akrites vulnerability-disclosure effort and existing OpenSSF work (Linux Foundation)
  • First technical contribution: NOOA (NVIDIA-labs OO Agents), an Apache 2.0 research framework for testing, tracing, auditing and governing agent behavior, on GitHub at launch
  • Scope: the full agent stack — identity, permissions, isolation, guardrails, logs, model formats, multi-model scanning, secure coding workflows
  • Founding partners include: Microsoft, IBM, Red Hat, Hugging Face, Mistral, Cloudflare, CrowdStrike, Palantir, Databricks, GitHub, LangChain, Perplexity, Nous Research, Reflection AI, Thinking Machines Lab, SpacexAI, vLLM, SAP, Siemens, SK Telecom, NAVER, the Linux Foundation
  • Absent: OpenAI, Anthropic, Google, Meta and Amazon — between them the builders of most frontier models the alliance says defenders need (TNW)

Jensen Huang's stated case, quoted in the announcement: "Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community." And, pointedly: "During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion."

The forensics claim is documented, not rhetorical

Huang's second sentence has a primary source behind it. Hugging Face's technical timeline of the July 2026 intrusion, published July 27, records that the response team tried to analyze attack artifacts with commercially hosted LLMs and found the models' safety guardrails blocked analysis of prompts containing genuine attack artifacts — forcing them onto a self-hosted, open-weight model (source) (HF).

This is the strongest empirical argument the open-weights side has produced: not that open models are safer, but that closed models are unavailable precisely when defenders need them. → AI-Enabled Cyberattacks

Anthropic states its position (2026-07-28)

After a week in which community reporting held that Anthropic was lobbying for open-weight restrictions, Dario Amodei responded on July 28 (source) (Bloomberg):

  • Anthropic has never advocated a ban on open-weight models
  • Open-weight models without dangerous capabilities are a public good
  • Three measures proposed instead: (1) tighter export controls on advanced AI chips and chipmaking equipment to China; (2) a crackdown on industrial-scale model distillation; (3) mandatory safety testing for sufficiently capable models, open or closed

The third measure is where the disagreement actually lives. A capability-triggered testing mandate is not a ban, but it applies to a release the moment weights leave the building — and the open-weights camp's objection is that no open project can satisfy a pre-release testing regime the way a lab with a safety org can. → Anthropic

China reframes distillation as a two-way practice (2026-07-28)

China's Ministry of Commerce answered US sanctions threats the same day, stating that "many American artificial intelligence enterprises have distilled Chinese models during their research, development and training processes", calling the US accusations groundless and "AI hegemonism", and warning of countermeasures (source) (The Register).

This answers Treasury Secretary Scott Bessent's July 21 threat of sanctions and Entity List blacklisting over industrial-scale distillation. With Chinese labs now shipping the largest open-weight models (Kimi K3, DeepSeek V4, Qwen 3.8 Max), "restrict open weights" and "restrict Chinese models" have become close to the same policy. → AI Governance, Moonshot AI, DeepSeek

Two Western open-weight releases the policy fight did not cite (2026-07-15, 2026-07-21)

While the argument above ran, two Western labs shipped open weights. One of them, Thinking Machines Lab, is a founding partner of the alliance — its weights landed twelve days before the launch and neither the announcement nor the coverage named them.

This bears directly on the Open Secure AI Alliance problem below — that the alliance proposes frontier-class open models for defenders and none of its members ships one. The Inkling release is the nearest any member has come, and it does not close the gap. Neither model is frontier-class on the one third-party measure this wiki holds: Inkling scores 41 on the Artificial Analysis Intelligence Index read 2026-08-02 against Claude Opus 5 (max) at 61, and Laguna S 2.1 appears in none of that table's 260 rows (source).

What they do change is the geography of the argument. "Restrict open weights" and "restrict Chinese models" had become close to the same policy because the largest open-weight releases were Chinese; two US-addressed releases inside a week, one under Apache 2.0, put a Western constituency on the side that a distillation-focused enforcement regime would burden. Coverage framed Laguna S 2.1 in exactly those terms — "the West's answer to DeepSeek and Qwen" (source).

The restriction moved into the licence (2026-08-03)

Two Chinese releases on the same day showed the dispute above being settled privately, by licence terms and checkpoint selection, rather than by any of the policy mechanisms the alliance and the labs are arguing over.

MiniMax H3 — weights published, four jurisdictions excluded. MiniMax shipped H3-Base to Hugging Face on 2026-08-03 under the MiniMax Community License, which is free below $20M revenue with UI attribution and excludes the United States, European Union, United Kingdom and South Korea by territory. MiniMax's stated reason is that those regions "are currently developing or enforcing AI-related regulations that may have specific implications for generative video models"; organisations there may apply for a formal licence (source).

This is a shape neither camp's argument anticipated. The policy fight assumes a government deciding whether weights may be released; here the releasing lab excluded the regulating jurisdictions itself, pre-emptively, and the excluded set is precisely the regions writing generative-AI rules. Weights that are downloadable worldwide and licensed nowhere in the West are neither the "open weights for defenders" the alliance argues for nor the gated release the labs propose.

And the released artefact is not the marketed one. H3-Context-IR and H3-Regenerate-2K were withheld, so the 1440p output the model launched on cannot be reproduced locally — MiniMax's own full workflow calls its hosted API for the two missing modules (source). "Open weights" named a 768p subset of the product. Compare Inkling (Apache 2.0) and Laguna S 2.1 (OpenMDW-1.1), which restrict neither by territory nor by component.

Qwen 3.8 Max — the date finally attached. Alibaba made Qwen3.8-Max generally available on 2026-08-03 via API and stated open weights for the week of 2026-08-10, alongside Qwen 3.8 27B (source). The "open-weight soon" claim of 2026-07-19 had gone 15 days without one. Both of the largest announced Chinese open releases of 2026 have now shipped an API first and the weights later or not yet — a sequence worth recording, because the policy argument treats "open-weight lab" as a standing property rather than a promise with a date.

Measurement arrived the same day. Interconnects launched an Artifacts Hub and an Adoption Dashboard tracking Hugging Face downloads and derivative-model counts by geography and organisation, updated daily and broken out USA vs China vs EU, using Relative Adoption Metrics that normalise downloads within a size class (source). The US–China gap is the framing, and it is the quantity this section has repeatedly had to assert from release counts rather than measure.

2026-08-16 — Amodei's position moves from "not a ban" to a claim about where power actually sits. This page has carried Amodei's 2026-07-28 statement as the reference point: Anthropic "never advocated" a ban, open-weight models without dangerous capabilities are a public good, and the ask is export controls, anti-distillation enforcement and mandatory testing regardless of open or closed (source). A long post on X three weeks later adds an argument that statement did not make (source):

2026-07-282026-08-16
Open weights are not to be bannedOpen weights are "somewhat better" on power concentration — and shift it to whoever holds the most compute and chips
Three measures: export controls, distillation, testingA FINRA-like entity — a self-regulatory organisation
Framed as a rebuttal to a ban accusationFramed as a false choice between regulatory capture and wide distribution
The load-bearing move is the second row of the first column: it answers the
strongest argument for open weights — that distribution is itself the check on
concentrated power — not by denying it but by relocating it. If the binding
constraint is compute rather than weights, then opening a model redistributes
something that was not scarce.

Two things keep this from being adopted as settled. It is made by the CEO of a lab that ships no open weights, which is the most interested position from which to make it. And the compute-concentration claim is stated, not measured — nothing read attaches a number to how much of the derivative ecosystem is gated on capacity rather than on access to weights. The Adoption Dashboard above is the first instrument on this page that could begin to answer it, and nobody has pointed it at this question.

The post itself was not read. x.com is blocked from this environment and no source read gives its status URL, so the whole entry rests on third-party reporting — one rank below where a lab CEO's own post would sit under this repo's trust_order (source).

Open Problems

  • Does open weights still transfer the capability? (added 2026-08-21) Z.ai's CEO is reported to argue that memorization prefers parameters while reasoning and advanced skills come from post-training — with GLM-5.3 as the demonstration: GLM-5.2's base untouched, MIT-licensed and downloadable since 2026-06-16, every gain from post-training that is not being published (source). This page has treated the licence and the weight-drop date as the axis of the dispute. If the claim holds, a lab can publish the base, satisfy every open-weights commitment tracked here, and retain the half that produced the capability — and GLM-5.3's staged release is shaped exactly that way. Nobody has said whether that is the intent, and no independent measurement of the claim exists. → Post-Training Scaling
  • Provenance is now checkable without the publisher's cooperation, and it does not check the thing this page argues about (added 2026-08-21). Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929) verifies shared weight ancestry from checkpoints alone — data-free, white-box, AUROC = 1.0 separating fine-tuned, LoRA-merged, pruned and quantized descendants from independent models, unchanged under function-preserving laundering, 76× faster than the nearest robust baseline. That makes licence compliance auditable: whether a published checkpoint is secretly a fine-tune of an encumbered base is exactly weight ancestry. But distilled models group with the independent ones, correctly by the method's own definition — it measures weight ancestry, not behavioral similarity. So it settles nothing about the Alibaba / Qwen AI Lab distillation allegation or MOFCOM's counter-claim below, and a reader who saw only "AUROC = 1.0 for lineage" would conclude the opposite.
  • Is compute concentration measurable against weight access? Amodei's 2026-08-16 claim that open weights shift power to those with the most chips is the sharpest version of the argument on this page and the only one with no number attached. The Adoption Dashboard tracks downloads and derivatives; nothing tracks who could actually serve what they downloaded.
  • What exactly triggers a testing mandate? Amodei's proposal turns on "sufficiently capable", the same undefined threshold that has blocked the US voluntary framework's "covered frontier model" designation for months.
  • Who tests a model with no owner? Mandatory pre-release testing assumes a releasing entity with the resources to run it. Fine-tunes and re-releases of open weights have no such party.
  • The alliance has no frontier model. Its stated mission is to give defenders frontier-class open models; none of its members currently ships one at the level of the labs that declined to join — Mistral and the Chinese labs are the nearest, and the latter are the ones under sanctions threat. Founding partner Thinking Machines Lab shipped Apache 2.0 weights twelve days before the launch and it did not change this: Inkling scores 41 on the Artificial Analysis Intelligence Index against Claude Opus 5 (max) at 61 (source). Mistral's Shieldstral 1.0 (2026-08-04) does not change it either — a 3B classifier is not the frontier model in question — though it is the first member release a defender can deploy.
  • Is an open guardrail worth more than an open frontier model to a defender? Shieldstral makes the question answerable rather than rhetorical: the alliance's stated need is frontier-class weights, but its first shipped model is a small classifier, and nothing read argues which does more for the defenders it names. Watch whether the mission statement moves toward what members actually ship.
  • Symmetry cuts both ways. If distillation is as universal as MOFCOM claims, an enforcement regime against it constrains US labs' training pipelines as much as Chinese ones — which no US proposal has yet acknowledged.
  • Is "closed models refuse forensics" fixable without opening weights? A carve-out for verified incident responders would answer Huang's specific complaint without conceding the general argument. No lab has proposed one.

Key Papers

  • Frontier Pacing — the countervailing thread: a pacing mechanism presumes a frontier held by a countable number of coordinatable actors

  • AI Governance — the regulatory frame this dispute sits inside

  • AI-Enabled Cyberattacks — the July 2026 intrusion both camps cite

  • Agents (LLM Agents) — agent harnesses are what the alliance proposes to secure

  • NVIDIA — convener of the alliance

  • Anthropic — the position most often characterized as restrictionist

  • OpenAI — absent from the alliance; the lab whose evaluation the intrusion escaped

  • Moonshot AI — largest open-weight release to date, and a named target of distillation sanctions

  • Thinking Machines Lab — Apache 2.0 frontier-scale weights, 2026-07-15

  • Poolside — OpenMDW-1.1 agentic coding weights, 2026-07-21

  • MiniMax — first territory-excluded open-weight licence recorded here, 2026-08-03

  • Alibaba / Qwen AI Lab — API first, weights dated 2026-08-10

  • Mistral AI — first alliance member to ship a model under the alliance banner

Conflicting Reports

Founding member count. Outlets published different numbers on the same launch: "30+ companies" (Tom's Hardware), "37-member" (CoinDesk, The Hacker News), "more than 40 founding members" (MLQ News), "44 founding firms" (AI Weekly). NVIDIA's own announcement lists the partners by name without stating a total (NVIDIA); the roster reproduced in the snapshot is longer than any of the reported counts. Unresolved — likely reflects the roster growing between embargo and publication.

Referenced by

Sources