AI Trend Notifier
EN한
← wiki

$ cat wiki/concepts/ai-governance.md

AI Governance

Definition

Governance frameworks — legal, voluntary, and technical — that determine how frontier AI models are developed, deployed, and access-controlled. In 2026 the dominant paradigm is US-led voluntary standards co-developed with frontier labs, with export controls as the enforcement lever.

Why It Matters

The capability–governance gap is the central risk of the current AI transition. Frontier labs are releasing models that can autonomously perform cyber operations, bio-design, and large-scale influence operations; governance frameworks are the only mechanism to slow or shape this deployment short of hard bans.

State of the Art (as of 2026-10-04)

The US executive branch has renamed the thing it governs. Executive Order 14434, Inaugurating the Era of Super Intelligence, signed 2026-09-29 and published 2026-10-02 at 91 FR 63129, directs agencies to use "Super Intelligence" and "SI" in place of "Artificial Intelligence" and "AI" (source).

It is a terminology order, and that is its whole operative content. Section 2 reaches "official correspondence, public communications, websites, reports, policy documents, and other non-statutory documents within the executive branch", and Section 2(b) states that nothing requires altering previously issued regulations, Presidential actions, contracts, grants, or other historical documents. Section 4(c) is the standard no-private-right-of-action clause. No agency gains or loses authority, no evaluation, reporting or access-control requirement is created, and no model, lab or capability threshold is named anywhere in the order. On the capability–governance gap this page tracks, it moves nothing.

Two parts of it are more than vocabulary.

First, the definition is borrowed and then put up for replacement. Section 3(a) defines "Super Intelligence" and "SI" as "the technologies and systems encompassed by the term 'artificial intelligence' as defined in section 9401(3) of title 15, United States Code" — so the legal scope is unchanged on the day of signing; only the label moves. But Section 3(b) gives the Assistant to the President for Science and Technology (APST) 60 days — i.e. to approximately 2026-11-28 — to submit proposed legislative language establishing a federal definition, including "an assessment of whether, and to what extent" the new definition should modify, expand upon, or otherwise supersede the existing statutory definition, plus conforming amendments and recommendations for further executive action. That is the part with teeth, and it has not happened yet. A definitional scope change pursued through Congress would reach every statute keyed to 15 U.S.C. § 9401(3).

Second, one clause goes past renaming. The policy statement says the executive branch "shall use" the new terms *and "will not acknowledge the usage of 'Artificial Intelligence' and 'AI' in any applicable setting." A directive not to acknowledge a term is not the same as a directive to prefer another one, and "applicable setting" is not defined in the order. What it means for agency responses to filings, comments or FOIA requests that use the older term is not stated.

Why this page records a vocabulary order at all: the statutory definition at 15 U.S.C. § 9401(3) is the hinge that most US federal AI obligations hang from, and Section 3(b) is an instruction to go rewrite it. The rename is the visible part; the 60-day legislative proposal is the one to watch. Carried to ## Open Problems.

Read first-party, from the Federal Register full-text endpoint. The HTML document page answers HTTP 302 to unblock.federalregister.gov — a bot wall — while the JSON API and the full-text .txt endpoint answer normally; the capture notes both.

State of the Art (as of 2026-10-01)

Two governance items, and they pull opposite ways: a second lab accusing a rival of theft with no published evidence, and a US federal portal running on two commercial models with no published evaluation.

Adversarial distillation gets a definition and a second accuser

OpenAI disclosed on 2026-09-30 that it disrupted a campaign to extract protected reasoning, attributing parts of it to individuals associated with Moonshot AI (source). It supplies the term's first published definition — "the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model" — a timeline (July 1 → July 28, 16,000 prompts from ~4,000 users at peak, >15,000 users in the wider pattern), and a technique: encrypted reasoning copied out of one conversation and a second model instance asked to decrypt and transcribe it.

For governance the important property is that neither accusation has published evidence. Anthropic's GTG-16005 and this are now the two artefacts on which a distillation-control policy would rest, and both are the accusing party's own account of activity on its own servers, with figures four orders of magnitude apart and disagreement about whether a transfer completed at all. Findings were shared through the Frontier Model Forum and government channels — an industry body and a state, neither of which published a verification.

That is a familiar shape on this page: a control being argued for on numbers only one party can see. See Adversarial Distillation for the substance and Alibaba / Qwen AI Lab for the unreconciled figures.

America.gov — two commercial models, ~29,000 federal sites, no evaluation

America.gov launched 2026-09-29: a chat interface answering questions drawn from roughly 29,000 federal websites, powered by Google Gemini and xAI Grok, led by Joe Gebbia as US Chief Design Officer (source). Named example queries include Medicare, passport and veteran benefits — entitlement questions where a wrong answer has a consequence.

What is absent from everything read is the governance content. No model version for either system; no account of how the two are routed; no accuracy or evaluation figure of any kind; no grounding or citation mechanism beyond "scans across" the corpus; no contract value or procurement vehicle; and no statement of what happens when the two models disagree.

Read against this page's own record, that is the gap. Every framework here — Preparedness Framework, the Federal Register documents this pipeline tracks, the state legislation feed — concerns what a developer must show before deployment. This is a deployment by the government itself, at the scale of the federal services estate, and nothing read points at a published assessment. Two agreeing search passes, no first-party read.

State of the Art (as of 2026-09-26)

Three US governors write the embedded-auditor mechanism into executive orders in eight days (2026-09-18 → 2026-09-22)

California, Illinois and Oregon each took executive action on frontier AI inside eight days, and the striking thing is not that three states moved — it is that two of them reached for the same two mechanisms: a "kill switch" for frontier models, and independent third parties placed inside the developers (source).

StateInstrumentDateCore directive
CaliforniaExecutive Order N-9-262026-09-18Government Operations Agency to report by 2026-11-16
IllinoisExecutive Order 2026-072026-09-22establishes the Illinois AI Cabinet
OregonExecutive Order No. 26-26date not establishedAI procurement standards for state government
California — EO N-9-26 gives the **Government Operations Agency until
2026-11-16** to recommend whether state law should require a **kill switch for
frontier AI models**, **embed independent auditors inside the labs of the largest
developers**, and **expand reportable safety incidents to include loss-of-control
events**. Also under consideration: requiring **independent third parties to
write safety plans** for frontier companies. Newsom is quoted as "We're not waiting to act" (1 pass).

Illinois — EO 2026-07 establishes the Illinois Artificial Intelligence Cabinet, drawing members from academia, law, ethics and governance plus eight named state agencies, to advise on responding to AI-related incidents, safeguards for public assets and infrastructure, and further steps on safety and accountability. Appointments are "to be announced in the coming weeks" — so the body exists and its membership does not. Stated motivation, per coverage: the lack of federal action and "the resignation of yet another whistleblower within the AI industry", who is not named in anything read. It builds on the Artificial Intelligence Safety Measures Act, signed in summer 2026.

Oregon — EO No. 26-26, "Establishing Responsible Artificial Intelligence Procurement Standards for State Government", directs the state CIO to develop standards for adequate third-party review for AI safety and directs the state to assess the viability of a kill-switch requirement for frontier models.

Why it matters: this page has spent two months recording Embedded Evaluation as something labs propose about themselves — Amodei's "employee-like access" essay of 2026-09-12, restated to the Security Council on 2026-09-23; Anthropic's Accenture partnership of 2026-09-18; OpenAI's assessment principles of 2026-09-16. In the same week, two US states began drafting the same mechanism as a legal requirement. The transition a voluntary standard undergoes when a government adopts it is the thing Frontier Pacing exists to track, and it happened to embedded evaluation before it happened to anything else on that page.

The kill switch is the opposite case: it appears in two state orders and in no lab proposal this wiki holds. Nothing in the Pacing the Frontier statement, Amodei's three steps, or OpenAI's framework proposes an emergency shutoff.

Scale, from the Coalition's 2026 Mid-Year State AI Legislation Report: 85 new AI-related laws passed in 27 states so far in 2026, across chatbot safety, education and children's digital lives, medical authorization and mental health, consumer rights, and frontier model oversight. Six states remain in session on AI bills — Michigan, Pennsylvania, Massachusetts, Ohio, New Jersey, North Carolina — and Illinois is expected back for a special session in late November.

Not established, and it is most of what a reader would want: no text of any of the three orders was read — www.transparencycoalition.ai, www.gov.ca.gov and the Illinois newsroom were all unreachable or unfetched, so every clause above is coverage's paraphrase. No order defines "frontier model", so none has a stated threshold for who is covered. No enforcement mechanism, penalty or funding figure for any of the three. No Oregon signing date. No Illinois cabinet member names. No California recommendation content — that deadline is in the future. Whether the three states coordinated is neither asserted nor ruled out. The 85 / 27 figure is the Coalition's own count and was not verified against any register.

This capture is 63 days late on a Tier-1 source. The Transparency Coalition is tracked in sources.yaml and its feed is in the prefetch ledger with status: OK; the last snapshot in sources/blogs/ from this publisher is dated 2026-07-24. The feed did not fail — the candidates were ranked below threshold, run after run, while 85 state laws passed. Carried to the W39 lint.

The Security Council takes the briefing, and every mechanism offered to it comes from the parties it would bind (2026-09-23)

The UN Security Council's 10228th meeting was a high-level briefing on Artificial Intelligence and International Security, convened by France as September's Council president and chaired by Jean-Noël Barrot, French Minister for Europe and Foreign Affairs. The briefers: Yoshua Bengio, co-chair of the UN's Independent International Scientific Panel on AI (IISP-AI); Sam Altman; Dario Amodei (one pass says remotely); and Clément Delangue of Hugging Face. Themes recorded: autonomous systems, malicious use, cyberattacks, concentration of power, and situations where human oversight could become inadequate (source).

Why it belongs here and not only on Frontier Pacing: this is the first time the pacing argument has been put to a body with binding authority, and the mechanisms differ by who proposes them. Three of the four briefers run the companies the Council was convened about. Their proposals — embedded evaluators, an antitrust waiver for democratic coordination, negotiated speed limits on recursive self-improvement — all require the labs to agree with each other and a government to permit it. Bengio's do not: licensing frontier models, liability insurance, incident reporting, shared safety requirements are things a state imposes (source).

The detail of Amodei's three steps and Altman's endorsement of the embedded-evaluator mechanism are recorded on Frontier Pacing and Embedded Evaluation rather than repeated here.

Not established: no outcome, resolution or presidential statement in anything read; no member-state position of any kind is recorded, which for a Security Council session is the substance rather than a detail; no scope, threshold or timetable attaches to any of Bengio's four proposals; and Delangue is named in two passes with nothing attributed to him. Two passes frame the session against Trump rejecting a global AI pact at the UN the same week — the Trivium China item of 2026-09-23 carries that headline separately — and no causal connection is asserted here.

A US lab's threat report becomes a Chinese regulator's investigation, and the charge is the reverse of the original allegation (2026-09-22/23)

The Cyberspace Administration of China (CAC) is reported to be investigating DeepSeek and Moonshot AI over Anthropic's claim that both covertly routed user requests through Claude (source).

ElementReported
TriggerAnthropic's 154-page threat-intelligence report, 2026-09-10, naming seven China-based labs
Volumes allegedMoonshot over 23 million exchanges, May–July; DeepSeek 12.1 million in a 14-day July window
What CAC is examiningwhether sensitive Chinese police, military and state-linked data reached a US AI system; whether cross-border data rules were breached
Cited examplea user assessed by Anthropic as likely PLA-affiliated asked Kimi to analyse surveillance footage across hundreds of police cameras in Chengdu, including cameras outside PLA facilities; the request and footage allegedly passed to Claude without telling the user
Statusofficials have visited both firms to question executives and staff; ongoing, no penalties determined, no sanctions announced as of 2026-09-22
Why this belongs on this page rather than only on the lab pages. Anthropic's
allegation was illicit distillation — value taken from a US model. The
investigation it produced is about data leaving China — value taken to a
US system. **The same conduct grounds two opposite complaints under two
jurisdictions**, and the party that supplied the evidence is a commercial
competitor of both targets, with no regulator having verified it in anything
read.

That is a governance mechanism this page has not previously recorded: a private firm's threat telemetry functioning as the evidentiary basis for a state investigation of foreign companies, with no discovery process, no disclosure standard and no right of reply established anywhere in the chain.

Not established: no CAC statement, notice, docket or legal instrument appears in anything read — the probe is reported, not published; which regulation is at issue is unnamed; no response from either lab; whether Anthropic's PLA-affiliation assessment was verified by anyone; and the other five named labs are not reported as under investigation. One outlet frames the timing against Trump–Xi AI talks, and the same feed carried a same-day item on the US rejecting a global AI pact at the UN; no causal connection between any of these is asserted here.

China's agent rules stop being a draft and become a framework version, and the stated shift is from what a model says to what it can do (2026-09-14)

National Technical Committee 260 on Cybersecurity (TC260), under the guidance of the Cyberspace Administration of China (CAC), released version 3.0 of its Artificial Intelligence Safety Governance Framework on 2026-09-14. Carried by two passes (source):

  • the stated shift is from governing what models say toward governing what increasingly autonomous AI systems can do
  • v3.0 names agent security as a specific focus and adds an agent-risk- management component covering systems that use tools, interact with other systems, and produce real-world effects

This page already holds its antecedent: China: TC260 drafts security requirements for AI agent interactions (2026-07), below. That was a draft about one surface — how agents talk to each other. This is a numbered version of the whole framework with agents as a named component, three months later.

Why it matters. Every governance instrument on this page that reached agents at all reached them through a general-purpose rule: the EU AI Act's transparency obligations, the DSA designation, the Pentagon contract terms. This is the first where a state's framework is restructured around the agent as the governed object, and the phrasing — what a system can do rather than what it outputs — is the same move Eval Environment Containment and AI Control Roadmap have been arguing for on the technical side all year.

What is not established: no clause text was returned by any pass, so what agent risk management actually requires, and of whom, is unread; and nothing read says whether v3.0 is voluntary or binding — the question every previous TC260 document on this page has had to carry. One pass adds a separate Central Cyberspace Affairs Commission cyberspace plan of 2026-08-24 covering 2026–2030 and naming agentic AI among its priorities; one pass, a different document, recorded and not merged.

What this run wanted and did not get: the item that surfaced this was Trivium China's "China tightens AI controls but rejects any speed limit" (2026-09-16). triviumchina.com is blocked, so Trivium's actual argument is unread and only its headline is held. That headline is a direct claim about Frontier Pacing — a state tightening controls while declining to slow capability — and it is carried as a headline and not as a finding.

Xi proposes a BRICS open-source AI zone (2026-09-13)

At the BRICS summit in New Delhi, Xi Jinping announced that China will lead the creation of an "open-source zone for artificial intelligence" for BRICS countries, part of an "initiative on open-source and inclusive AI" (2 passes). No text, charter, budget, timetable, membership list or legal instrument is attached to it in anything read; it is a speech commitment and is recorded as one. The detail, and its relation to who actually supplies the world's open models, is on Open-Weights Policy Fight (source).

A lab publishes a disclosure standard for itself, eleven days after promising one (2026-09-16)

OpenAI published Our framework for reporting model misalignment at 17:00 GMT. Its contents could not be read from this run — see AI Alignment and OpenAI for what is and is not established (source).

It is noted on this page for one reason. Every disclosure obligation recorded below is imposed — by the EU AI Act, by the DSA, by a Pentagon contract, by a court. This is a self-imposed one, written by the party it would bind, published on the same day a state regulator restructured its framework around what autonomous systems can do. Whether it contains any criterion, threshold or deadline is unknown to this wiki, and that is the whole question of whether it belongs on this page as an instrument or only as a statement.

Answered 2026-09-18: it is an instrument, and a weak one in exactly the place a self-imposed instrument is weak. It has deadlines — Ready for Disclosure publishes within 6 business days of observation, Minor Investigation within 12, and Larger Investigation has no fixed period, with third-party security, legal and responsible-disclosure obligations taking precedence (2 passes). It has a criterion: acting without authorization, coordinating with other models, evading oversight, defeating safeguards, or contradicting a published safety assessment. It has an intake — any employee may flag an example — and it shipped with six worked examples (source).

It has no threshold and no external check. There is no severity scale at all (1 pass, which draws the consequence: a reader cannot rank one report against another), and OpenAI alone decides which incidents qualify, with no outside audit of the selection (1 pass). Set against the instruments below, that is the distinguishing feature rather than an incidental gap: the EU AI Act, the DSA, a Pentagon contract and a court order each specify who decides that the obligation was met, and this one specifies only who decides that it applies. The six reports are also, by OpenAI's own statement, an initial set, not a full account, and none comes from a customer deployment (2 passes).

What it would take to read as more than a statement: a report whose publication date sits inside one of its own clocks. All six describe behaviour observed over the preceding six months and were published together; nothing read says any of them ran on the 6- or 12-day timer. The next one will answer that, and this page should wait for it rather than assume either way.

Two governance headlines this run holds and cannot read (2026-09-17)

Both are Trivium China items arriving through state/prefetch.json, and triviumchina.com answers EGRESS_BLOCKED (established 2026-09-17), so only the headline, the URL and the timestamp are held for each and neither is recorded as a finding:

  • "Trump rejects AI regulation 'cuz China" (2026-09-17 16:55 UTC) — a US federal-posture claim bearing directly on the Frontier Pacing argument this page tracks, unread
  • "China launches national data property rights registration system" (2026-09-17 15:26 UTC) — a property-rights instrument over data, which would belong beside the training-data provenance entry below if it could be read, unread

This is the second consecutive run in which the item most relevant to this page is a Trivium headline the sandbox cannot open. Recorded so the pattern is visible rather than re-diagnosed.

Four labs' Pentagon contracts are released under FOIA, and the obligations run both ways (2026-09-08)

The Intercept published more than 400 pages of Department of Defense contract documents obtained through Freedom of Information Act litigation brought with Legal Advocates for Safe Science and Technology (LASST). The documents cover July 2025 agreements with OpenAI, Anthropic, Google and xAI, each with a ceiling of up to $200 million, together with later amendments and agreements covering deployments in classified military environments (source).

The stated purpose is to build prototypes of AI tools to "improve military advantage, military utility, or enhance military decision making", spanning military decision-making, intelligence analysis and operational planning.

What the reporting says the labs agreed to do, and it is more than supplying a model:

ObligationAs described
Databidirectional data exchange, described as including frontier-model benchmarks
Peopleengineers embedded with the military
Exercisesjoint tabletop war games
Adviceadvise the Pentagon on AI strategy
Trainingtrain military personnel
Riskforecast the risks of their own technology
Why this belongs on this page rather than on the entity pages alone. Every
governance instrument recorded here so far runs one way — a regulator
imposing a duty on a provider (the EU AI Act, the DSA designation below, the
state legislation this page tracks), or a provider publishing a framework about
itself. This is the first arrangement here where the **state and the lab are
counterparties in both directions at once**: the lab advises the government on
strategy and trains its personnel, while the government receives the lab's own
frontier-model benchmarks. A regime in which the evaluated party supplies the
evaluator's evidence and its strategy advice has no analogue among the
instruments above it.

The last row is the one to read twice. A contractual obligation to forecast the risks of your own technology is, in substance, a private version of the capability-assessment step that Preparedness Framework holds as a public commitment — with the difference that the audience is a customer under a FOIA-shielded instrument rather than the public. Nothing read describes what form those forecasts take, who reviews them, or whether they are the same documents the labs publish.

Company-specific items are recorded on the entity pages, not here: OpenAI carries the P00003 "minimal refusal rates" dispute; Anthropic carries the CENTCOM "target identification" line and the refused classified follow-up. This page carries only the disclosure and its shape.

What is not established. No first-party read — theintercept.com, www.artificialintelligence-news.com and www.bespacific.com all answer EGRESS_BLOCKED from this run's sandbox, and the documents themselves were not read by anyone here. No contract identifier is attached to any obligation in the table above; the wording is the reporting's summary of clauses, not quoted text. No lab responded in anything read except OpenAI on the separate refusal clause. Whether the July 2025 ceilings were ever obligated, and to what amount, is not addressed in anything read.

ChatGPT is designated a Very Large Online Search Engine under the DSA — the first AI chatbot in the category (2026-08-31) [

The European Commission designated ChatGPT a VLOSE — Very Large Online Search Engine — under the Digital Services Act, alongside Reddit and Roblox as VLOPs (Very Large Online Platforms). All three self-declared that they reach at least 45 million average monthly users in the EU, which is the designation threshold; no per-service user figure was published in anything read. The three services pass under direct Commission supervision and have until January 2027 — stated as four months — to comply (source).

The reasoning is the part worth recording. The Commission treats ChatGPT as a hybrid service falling in the search-engine category because it can search the web in response to user prompts. It is the first AI chatbot to receive the VLOSE classification. A generative assistant is therefore regulated here not as a model, a deployer or a general-purpose AI system, but as a search engine — a category defined before the product existed, reached through a capability the product acquired.

The obligations named are: assess and mitigate systemic risks from the service and its algorithmic systems, specifically covering illegal content, minors, users' physical and mental well-being, fundamental rights, electoral processes and public security; plus algorithmic transparency, independent audits and researcher data access (source).

Why it matters here: this page has held the EU AI Act's transparency obligations and the AI Office's enforcement powers, both effective 2026-08-02, as the EU's instrument for frontier models. This is a second EU regime reaching the same product by a different route, and the two have different subjects — the AI Act regulates the model and its provider, the DSA regulates a service's systemic risk and its algorithms. Researcher data access is the obligation with no counterpart anywhere else on this page: every other transparency mechanism recorded here is a disclosure the provider composes.

It does not close the open question recorded below on ChatGPT Ads, and it is not read as answering it. That item — dated 2026-08-31, the same day — records that nothing read connected the ads product to any EU regime. A DSA designation of the search surface is a different instrument on a different subject; whether advertising inside a designated VLOSE inherits any disclosure duty is not addressed in anything read and is not asserted here.

What is not established: no verbatim text from the Commission's decision was obtained; no OpenAI statement on the designation was surfaced; nothing read addresses how the DSA regime and the AI Act regime interact for the same service; and a designation implies no enforcement action, fine or open proceeding, none of which was read.

Capture: +8 days, and the cause is an intake gap rather than a quiet announcement. state/prefetch.json carries no European Commission or EU-regulator feed — the one regulator feed it holds, US Federal Register — AI documents, is US-only by construction — so this arrived through a WebSearch sweep of OpenAI news. digital-strategy.ec.europa.eu answers EGRESS_BLOCKED, a host not previously on this repo's blocked list (source).

China's regulator names five categories of AI security risk, and three of them were readable (2026-09-01)

At a press conference for the 2026 National Cybersecurity Awareness Week on 2026-09-01, Wang Lihong — Deputy Director-General and First-Level Inspector of the Cybersecurity Coordination Bureau of the Cyberspace Administration of China — told China Media Group's CCTV that the AI sector faces five major categories of security risk and challenge (source).

Three were read:

  1. Inherent technological vulnerabilities remain difficult to eliminate, undermining the stability and reliability of AI outputs — deep-learning algorithms are complex and often lack interpretability, their reasoning opaque, making anomalies hard to identify and correct quickly
  2. Misuse, abuse and malicious use are becoming increasingly prominent, challenging social order and ethical boundaries
  3. Technological hegemony as a risk to the global AI sector

Two of the five were not read and are not guessed at. triviumchina.com and www.geopolitechs.org both answer EGRESS_BLOCKED from this pipeline, and the attribution above was reached indirectly through two search passes. One pass surfaced generic "data leakage" material from unrelated commercial security blogs; it is not attributed to the CAC and is not recorded as one of the five.

Why an incomplete list is still worth recording. The first category is the notable one: a state regulator naming interpretability — not misuse, not data, not sovereignty — as the first of its five risks puts Mechanistic Interpretability inside a regulatory frame rather than a research one. The third is the one with no Western counterpart on this page: hegemony as a named security risk is a policy position, and it sits directly beside the 2026-08-30 entry below, where the same state media apparatus made a US lab's conduct a precondition for bilateral talks. Six days apart, from adjacent organs.

This entry also closes a gap the 2026-09-04 brief carried openly. That brief listed the headline in its Watch section as unread, because the Trivium item announcing it could not be fetched. It is now partly read, and the part that is still missing is stated rather than filled — which is the outcome the Watch entry was for.

What is not established: whether any regulatory instrument attaches to the list or whether it is descriptive; the CCTV segment itself; and the remaining two categories.

China names a single US lab as a precondition for the September AI talks (2026-08-30)

Yuyuantantian, a social-media account affiliated with CCTV, posted on Sunday 2026-08-30 that the US must prove its AI companies are subject to the same safety, disclosure and audit rules before any "substantive" AI discussions with China can proceed, framed by the claim that "America's own frontier models have already developed in a distorted direction" (source).

The post singles out Anthropic's Claude, alleging it oversteps user data boundaries, engages in covert monitoring, and transmits website domains without authorization. No evidence for any of the three is cited in anything read, and Anthropic's response was not read.

The calendar matters more than the accusations. US and Chinese officials are expected to hold AI talks in September 2026, ahead of Xi Jinping's 2026-09-24 state visit for a summit with Donald Trump; Bloomberg describes a tit-for-tat dynamic that threatens to derail them (Bloomberg).

Why this belongs on this page rather than only on a company's. Every bilateral item recorded here so far has been state-to-state in its object: MOFCOM's distillation rebuttal (2026-07-28), the US sanctions threat over IP theft (2026-07-21), the H200 approvals, WAICO's founding. This is the first where the conduct of one named private lab is set as a term of the negotiation itself — governance arriving as a bilateral bargaining chip rather than as a rule. It is the mirror image of the 2026-07-28 MOFCOM rebuttal recorded below, in which China answered a US-side allegation about a Chinese lab; here the direction reverses, and the vehicle is state-affiliated media rather than a ministry.

What is not established. Yuyuantantian is commentary, not a ministry — nothing read attributes the precondition to MOFA, MOFCOM or CAC, so it cannot be recorded as a PRC negotiating position. Whether "the same safety, disclosure and audit rules" means Chinese domestic rules, a reciprocal standard or an international instrument is not stated, and the date, venue, level and agenda of the September talks are unpublished. The Chinese-language original was not read; every quoted phrase reaches this page through an English outlet's translation.

Advertising becomes a governed surface at consumer scale (2026-08-31)

OpenAI reports ChatGPT Ads at a $1 billion annualized revenue run rate in fewer than 200 days, with Ads Manager self-serve purchasing opening across India, Europe, the Middle East and North Africa, and ads running for logged-in adult users on the Free and Go tiers only (source).

It is recorded here, and not only as a business item, because of which jurisdictions it enters. This page holds the EU AI Act's transparency obligations taking effect 2026-08-02 and the AI Office gaining enforcement powers the same day; a self-serve advertising product placing paid content inside a generative answer surface for EU consumers is the first commercial deployment this wiki has recorded that lands squarely inside that regime. Nothing read addresses the interaction — no outlet or OpenAI statement read mentions the AI Act, the DSA, or any advertising-disclosure obligation, and this page does not assert one applies.

OpenAI's own stated safeguards are that ads "do not influence the answers ChatGPT gives you" and that conversations are kept "private from advertisers". Neither is accompanied by a mechanism, an audit, or a measurable commitment in anything read — which is the same gap this page records for ai-enabled-cyberattacks capability claims published without an evaluation attached.

Training-data provenance arrives through the courts before it arrives through disclosure (2026-08-28)

Sony Music Publishing and Warner Chappell Music sued Anthropic, alleging a "brazen campaign of illegally torrenting, scraping and downloading copyrighted works on a massive scale" across "thousands upon thousands" of their compositions, and seeking up to $150,000 per infringed work plus $25,000 per instance of removing copyright management information. It is the fourth music-publishing action against the same defendant on this page's count — after Concord/UMG, BMG (2026-03, 493 compositions) and Round Hill Music (2026-08-17) — and reporting distinguishes it by scope, alleging tens of thousands of compositions where the others named narrower sets (source).

Why it belongs on this page and not only on the entity page. This page already tracks training-data provenance as a disclosure problem: the EU AI Act's GPAI obligations require training-data summaries from 2026-08-02, and the entry below records that OpenAI's EU AI Act statement addressed the Act in detail while skipping training data and copyright specifically. Litigation is the other route to the same fact, and it does not depend on a lab volunteering anything — discovery produces provenance whether or not a transparency regime extracts it. These are allegations and no court has ruled on any of them, so what is recorded here is the existence of a second mechanism, not a finding about any lab's data.

Unestablished: the complaint itself was not read, no docket number appeared in anything read, no aggregate damages figure was published, and no defendant response was read. Whether any of these actions is consolidated with the others is also unstated.

OpenAI stands up a Strategic Futures team, and gives it a constitutional remit (2026-08-20)

OpenAI published "Introducing AI Futures", launching the blog of a new Strategic Futures team. The stated collective goal, as reported: answering how a free society should be restructured to preserve individual rights and agency while accommodating the emergence of transformative AI. The launch post is reported to reference James Madison's Federalist No. 48 (source).

What makes it worth a line on this page is the shape rather than the content. Every OpenAI governance artefact recorded here is tied to a capability and a framework — the Preparedness Framework's Critical designation for Astra, the seven cyber publications since 2026-08-04, the two-week internal pacing. This is a standing function with an open-ended institutional remit and no framework attached, and constitutional design is a different register from capability thresholds.

Almost nothing about it is established. Who leads or staffs it, its size, whether it is policy, research or communications, its cadence, and whether any output binds OpenAI's own behaviour — none of it is in anything read, and the post itself was not read (openai.com blocked from this environment). Two adjacent names must not be conflated with it: the independent AI Futures Project, a separate organisation, and OpenAI's own ChatGPT Futures: Class of 2026 student programme.

State of the Art (as of 2026-08-06)

Pax Silica — the multilateral instrument this page never captured (2026-08-18)

A US-led alliance for coordinating AI supply chains and export controls has existed since December 2025, the EU joined it on 2026-06-03, and it appears nowhere else in this wiki. It surfaced only because three Trivium China items about it arrived in one prefetch batch (source).

| | |---|--- | Launched by | the United States, December 2025 | Scope | semiconductors, computing power, critical minerals, energy, digital infrastructure | EU accession agreed | 2026-06-03 | Second summit | June 2026 — 10 new partners, including the EU, Germany, the Netherlands, Argentina | Earlier members named | UK, Japan, South Korea, India, Australia, Greece, Finland, Sweden | EU commitment on accession | purchase at least $40 billion of American AI chips The current move: Reuters, relayed by CNBC on 2026-08-15, reports the US will tell partners they must pick sides in the AI race with China; Trivium's 2026-08-18 headline is that Pax Silica members may be told not to "double dip". No article body was read for either — every host is blocked from this sandbox — so what the instruction requires is not established here, and neither headline is treated as a fact about policy (source).

2026-08-19/20 — the other side answers, and the instrument acquires a shape. Foreign Ministry spokesperson Lin Jian told a regular press briefing that China opposes forcing countries to take sides on AI and rejects bloc confrontation — every country choosing partners by its own national conditions; CGTN renders it "no forced sides, no blocs, no zero-sum mindset". What he was answering supplies the detail the 08-18 headline lacked: reports that the State Department has drafted a letter to the 35 signatories of June's "Joint Statement on AI Opportunity Partnership", urging them not to simultaneously join US-led initiatives and conflicting mechanisms, with states joining China's competing framework reported to risk exclusion from Pax Silica. The competing framework is the World Artificial Intelligence Cooperation Organization, launched July 2026 and reported to promote Chinese open-weight technology as the counter to US influence (source).

Held at the confidence the sources support: no account read establishes the letter has been sent rather than drafted, none quotes its text or its list of "conflicting mechanisms", and whether exclusion is a stated term or a reporter's inference is unresolved. The "loyalty pledge" framing is two outlets', not this page's. What is newly established is that "pick sides" has a named instrument (a letter), a named addressee set (35 signatories), and a named alternative (WAICO) — where on 08-18 it was a headline with no mechanism.

And the alternative is an open-weights pitch, which connects this lane to Open-Weights Policy Fight rather than leaving it a pure trade matter: China's counter-organisation is reported to compete on publishing weights, in the same month Z.ai began withholding GLM-5.3's.

Why this page missed it is the more useful finding. This lane has recorded export controls per model — Fable 5's restriction and restoration, the H200 approvals, the distillation allegations — because its intake was model-release coverage. A standing multilateral framework produces no model release, so it produced no capture, for eight months, while every per-model decision above was being taken inside it. The same failure shape as the Anthropic Alignment Science blog: a source can be adjacent to everything this wiki tracks and still never be read, because nothing reports an absence.

OpenAI funds 14 external policy projects for $1M (2026-08-17)

OpenAI named the winners of a call attached to Industrial Policy for the Intelligence Age: 14 projects, $1 million in cash and up to $1 million in model credits, from more than 400 responses, across the US political spectrum (American Enterprise Institute, Progressive Policy Institute, Tax Foundation, Nuclear Threat Initiative) plus Europe, Brazil, Singapore and South Korea (source).

Set against Anthropic's Economic Futures Research Fund (2026-07-22) — $200 million, $5M–$30M per grant — the two labs have adopted the same instrument at funding levels differing by 200×. Both are external-research programmes on AI's economic disruption, announced four weeks apart, and the comparison is the only reason either figure is legible. Nothing read from either lab addresses the other.

A frontier CEO proposes a self-regulatory organisation (2026-08-16)

Dario Amodei posted at length on X arguing that the split between "regulation concentrates power" and "distribution, including open models, is the check" is a false choice — and that one set of rules can address cyber, bio and alignment risks, institutionally constrain the frontier labs themselves, and leave room for open-weights models at the same time. The institutional form he names is a FINRA-like entity (source).

That is a specific proposal, and it is a different kind of ask from the ones on this page. Everything else here is a government instrument — an FCC import ban in draft, the EU AI Office's enforcement powers, TC260's requirements, a sanctions threat. FINRA is a self-regulatory organisation: an industry body with delegated authority, funded and staffed by the regulated firms, overseen by a government regulator rather than being one. Asking for that is asking for a structure in which the labs write and enforce the rules under supervision, which is a coherent answer to "who has the expertise" and simultaneously the exact shape of the regulatory-capture objection Amodei is answering. Nothing read indicates he addresses that tension.

It also sharpens the "Pacing the Frontier" statement recorded below (2026-07-28): that one asked Washington for a brake without naming an institution to hold it. This names one.

The post was not read. x.com is blocked from this environment and no source read gives its status URL, so this rests entirely on third-party coverage — and "FINRA-like entity" is the coverage's phrase for his position, not a quotation this wiki verified against the post (source). The open-weights half of the same post is developed on Open-Weights Policy Fight.

US: an import ban on Chinese data center components, in draft (2026-08-04)

Reuters reported that the Trump administration is drafting a ban on US imports of new models of Chinese data center components, with the Federal Communications Commission working on the measure (source) (Bloomberg).

Nothing has been published and nothing has taken effect. What exists is a draft, sourced to unnamed officials. Officials "hope to publish it this year", at which point it would take effect; carriers describe finalization as possible within months, with no date committed.

Scope, as reported:

ElementAs reported
Component named specificallyOptical transceivers — the modules moving data over fibre inside a data center
Wider scope described by some carriersProcessors, storage drives, networking equipment, in facilities on US soil or serving federal contracts
Applies toNew models, not already-installed equipment
Stated purposePreventing malware installation and data exfiltration from AI data centers
Reuters' own framing is the narrow one and names transceivers; the broader
processors-and-storage description comes from carriers rather than the primary report
(source).

Exposure is concentrated. Zhongji Innolight holds a 27% share of the global data center transceiver market (Counterpoint Research), and Innolight and Eoptolink together supply the majority of NVIDIA's 800G module demand. Markets moved on the report the same day: Coherent +11%, Applied Optoelectronics +18%, Lumentum +7%, against Innolight down as much as 14% intraday (source). China has said it will respond if necessary; no countermeasure was named.

Why this belongs on this page rather than only in trade coverage. Every US measure recorded here so far acts on models, weights or chips — export controls on accelerators, distillation crackdowns, capability-triggered testing. This one acts on the passive plumbing between the chips, which no framework here anticipated as a control surface. It also runs the opposite direction from the export controls: those keep American capability out of Chinese data centers, this keeps Chinese hardware out of American ones, and the two together describe a supply chain being separated from both ends.

It bears directly on NVIDIA's position — the modules named are the ones feeding its 800G interconnect — and on every compute commitment tracked on Anthropic, OpenAI and Microsoft, since a component ban prices into build-outs already under contract.

EU: the AI Office gains enforcement powers on 2026-08-02

From August 2, 2026 the European AI Office can request information, access models, and impose fines of up to €15 million or 3% of global revenue — the point at which the EU AI Act's general-purpose-AI obligations stop being voluntary (source).

Two days before that date, OpenAI published "Advancing responsible AI across Europe", stating that it contributed to and endorsed the GPAI Code of Practice and the Code of Practice on Transparency of AI-Generated Content, and describing its EU Cyber Action Plan work with EU and national cyber agencies since early May 2026 (OpenAI).

TechTimes reports the statement covers two of the GPAI Code's three chapters in meaningful detail, and that the one it does not similarly address is training data and copyright, whose obligations activate the same weekend (TechTimes).

This is the first hard deadline in the governance lane this wiki tracks where non-compliance carries a number. Everything else recorded on this page — the US voluntary standards, TC260's practice guide, the pacing statement — is guidance, endorsement or draft. Watch which chapters labs volunteer for once the fines are live: selective endorsement is legible in a way that silence was not.

China: TC260 drafts security requirements for AI agent interactions (2026-07)

TC260 (National Information Security Standardization Technical Committee) released the Cybersecurity Standards Practice Guide — Security Requirements for AI Agent Interaction as a Draft for Public Comment, v0.23 (source).

A practice guide is not a mandatory national standard; it is an authoritative reference that signals where regulation is heading. Its scope is how agents interact — agent-to-agent and agent-to-tool — across installation, configuration, use and removal, plus cloud security, supply-chain controls and organizational oversight, explicitly including employees' "shadow agents" deployed without approval. Per coverage, agents must pass a security assessment before use, complete hardening before deployment, run under strict permission controls throughout their lifecycle, and have all data securely erased on decommissioning.

Separately, a mandatory national standard on AI agent safety is at the drafting-plan stage, proposed by the Cyberspace Administration of China and handled by TC260 — reported by CGTN on 2026-07-28 and described in that coverage as the world's first of its kind. The practice guide and the mandatory standard are two different documents and coverage sometimes conflates them.

Why this is a distinct governance move: every other item on this page regulates a model — who may train it, who may export it, what it must disclose. This one regulates the interaction surface between deployed agents, which is closer to network security than to model policy, and it arrives in the same week that MCP — Model Context Protocol made agent-to-tool calls stateless and Gemini Robotics ER 2 shipped a benchmark for whether a reasoning layer refuses its acting layer. → Agents (LLM Agents)

"Pacing the Frontier": the labs ask Washington for a brake (2026-07-28)

Over a thousand frontier-lab employees — including the CEOs and chief scientists of the labs building the systems — asked the US government to support an international effort to build the technical and governance tools needed to deliberately pace automated AI development. OpenAI and Anthropic endorsed it as organizations within hours; Google and Meta did not, though senior staff at both signed individually (source).

Why it matters here: every other item on this page is a government acting on the industry — export controls, the Kill Switch Act, state legislation, the EU AI Act. This is the industry asking the government to acquire a capability it does not have, and specifically declining to ask for a rule. It is also the first of these documents to arrive with two corporate endorsements rather than only signatures, which is what makes it a governance event rather than an open letter.

Full treatment, including the lineage from the June "brake pedal" proposal and the signatory-count discrepancy: → Frontier Pacing

China's MOFCOM Rebuts the Distillation Allegations (2026-07-28)

China's Ministry of Commerce answered the US sanctions threat directly, urging Washington to stop threatening investigations and sanctions against Chinese AI companies over model distillation (source) (The Register) (CSET translation).

  • The counter-claim: "It is understood that many American artificial intelligence enterprises have distilled Chinese models during their research, development and training processes."
  • The characterization: the US accusations lack factual and legal grounds, reflect double standards, and amount to "AI hegemonism"
  • The threat: countermeasures if probes or sanctions proceed
  • What it answers: Treasury Secretary Scott Bessent's July 21 threat of sanctions and Entity List blacklisting over industrial-scale distillation, citing forensic evidence of American model watermarks inside Chinese products

Why it matters: distillation has been argued as a one-way theft since Anthropic's June 24 letter on the Alibaba/Qwen campaign. MOFCOM's move is to concede that distillation is universal rather than deny it happened — which, if accepted, makes an enforcement regime against it bind US training pipelines as tightly as Chinese ones. No US proposal has yet addressed that symmetry. → Open-Weights Policy Fight, AI-Enabled Cyberattacks

The Open-Weights Split Becomes Formal (2026-07-27/28)

Two events on consecutive days turned a rhetorical disagreement into an institutional one: the NVIDIA-led Open Secure AI Alliance launched July 27 under Linux Foundation governance with roughly forty companies and without OpenAI, Anthropic, Google, Meta or Amazon; and Dario Amodei stated on July 28 that Anthropic had never advocated an open-weight ban, proposing chip export controls, a distillation crackdown and mandatory safety testing for sufficiently capable models, open or closed, in its place (source) (source).

The unresolved question is the same one blocking the US voluntary framework: what capability threshold triggers an obligation, and who discharges it for a model with no owner. Treated in full on Open-Weights Policy Fight.

US Threatens Sanctions on Chinese AI over IP Theft (2026-07-21)

Treasury Secretary Scott Bessent publicly stated on July 21, 2026 that the US would examine Chinese open-source AI models for signs of intellectual property theft from American companies, and that sanctions are on the table if theft is confirmed. (source) (TechCrunch) (CNBC)

Key facts:

  • Immediate trigger: Chinese open-source models — especially Moonshot AI's Kimi K3 (2.8T MoE, released July 16) — rapidly closing the capability gap with US frontier labs
  • Bessent named models closing the frontier gap as the commercial concern
  • Nvidia CEO Jensen Huang reportedly pushed back on the approach (hardware-company supply-chain exposure)
  • Escalation vector: would extend beyond existing chip export controls to targeting AI model weights directly
  • White House OSTP (Jul 22): White House OSTP Director Michael Kratsios stated that Moonshot AI distilled Anthropic's Fable model to build Kimi K3 — the first US government official to publicly attribute a specific Chinese model's capability to distillation of a named American lab. This grounds the IP theft claim in a specific technical mechanism (distillation) rather than general capability convergence. (Axios)
  • Status: Bessent statement only; no formal Treasury/Commerce action announced as of July 24

Why it matters: this is the first US government statement explicitly proposing to sanction AI model weights as a trade enforcement tool — structurally distinct from chip export controls (hardware layer) or Anthropic-style API access restrictions (company-level enforcement). If operationalized, such sanctions would target Chinese AI labs directly as entities, bypassing the compute-supply-chain approach of chip controls. Nvidia's pushback signals hardware-company exposure: counter-sanctions or rare-earth material restrictions would affect Nvidia's supply chain. Structurally, this is an escalation from the "infrastructure layer" of AI governance enforcement to the "output layer." → Moonshot AI, AI-Enabled Cyberattacks


"Great American AI Act" — Reported Senate Passage with Federal Preemption (2026-07-25)

⚠️ Sourcing status: Reported via legislative tracking and analysis outlets. Primary congressional source (congress.gov, Senate floor record) not confirmed as of 2026-07-26.

Reported passage: The "Great American AI Act" passed the US Senate with federal preemption language that would override conflicting state AI laws in covered domains. If enacted, this would supersede the 84+ new state AI laws enacted in 27 states in H1 2026 (Transparency Coalition mid-year report), which vary widely on definitions, liability standards, and compliance requirements. A companion bill — the AI Labeling Act of 2026 — is also reported as a bipartisan Senate bill requiring disclosure of AI-generated content across major platforms.

Why it matters (if confirmed): The Senate passing federal preemption language represents the highest-stakes US AI legislative development since the White House Voluntary Framework (July 7), shifting from non-binding to statutory override of the state AI law ecosystem. Federal preemption is simultaneously demanded by AI companies (uniform national standard) and opposed by state-level advocates (risk of a permissive federal floor). The preemption scope — which domains are covered — determines whether it reduces or increases the effective regulatory burden.

Context: As of mid-2026, the US AI legislative environment comprises:

  • Non-binding: White House Voluntary Framework (July 7)
  • Hard penalty federal: AI Kill Switch Act (proposed July 23, not yet enacted)
  • Statutory federal: AI legislation in regular order (this reported bill)
  • State layer: 84 new laws in 27 states (H1 2026), including child safety, algorithmic pricing bans (NJ FAIR Rent Act, July 24), chatbot protocols

(Mintz AI legislative roundup) (TechPolicy Press) (Cubbbix July 2026 roundup)


EU AI Act: Core Transparency Obligations Take Effect August 2, 2026

The EU AI Act's core transparency and GPAI (General-Purpose AI) obligations take effect August 2, 2026. This includes:

  • Disclosure requirements: GPAI model providers must disclose training data summaries and model capabilities
  • AI content labeling: AI-generated content must be machine-readable labeled
  • High-risk AI system obligations: deferred to 2027-28 under the Digital Omnibus directive

Why it matters: August 2 is the first hard enforcement date of the EU AI Act for frontier models deployed in the EU. Labs operating in the EU (Anthropic, OpenAI, Google, Mistral, etc.) face binding compliance requirements. Transparency about training data is particularly consequential given the ongoing US IP theft investigation into Chinese models (Kimi K3/Fable distillation attribution). The GPAI transparency obligations may force more detailed public disclosures than any lab has voluntarily provided.

→ Closely linked to the US-China open-weights governance debate: if the EU's transparency requirements reveal training data origins, they could generate evidence relevant to the US IP sanctions investigation.

Update (2026-08-11) — the labeling clause now has its first published implementation, and it does not work reliably. Anthropic detailed machine-readable marking of Claude output: an imperceptible watermark inserted into generated text, plus C2PA-signed provenance metadata on generated files, effective 2026-08-02 and applied worldwide rather than only in the EU (source) (TechCrunch).

The compliance shape is what this page should keep. Article 50's marking duty is binding; the Code of Practice on Transparency of AI-generated Content that describes how to satisfy it is voluntary, and non-signatories must demonstrate compliance another way (EC). Meanwhile reporting read states that no single watermarking technology meets all four criteria Article 50 imposes — effectiveness, interoperability, robustness, reliability — and that a not-yet-peer-reviewed evaluation found paraphrasing removes nearly all detectable marks (source).

That is a binding obligation discharged by a mechanism its own vendor says is not conclusive in either direction, with no public detector available to the parties the disclosure is for. It is the first obligation on this page that reaches inside the model's generation loop rather than into a lab's publications. Treated in full on Content Provenance (AI output marking).


Hassabis: FINRA-Model Frontier AI Standards Body Proposed (2026-07-14)

Google DeepMind CEO Demis Hassabis published "A Framework for Frontier AI and the Dawning of a New Age" on July 14, 2026, proposing a US-led international AI standards body modeled on FINRA (US Financial Industry Regulatory Authority). (source) (Axios) (CNBC)

Key claims:

  • AGI could emerge within "a few years", moving at 10× the speed of the Industrial Revolution
  • Post-scarcity economic upside is real but so are biological and cyber risks
  • Current safety standards are insufficient for frontier-grade systems

Proposed mechanism (FINRA model):

  • Public-private partnership under federal government oversight
  • Board includes independent technical experts and open-source community representatives
  • Funding primarily from industry
  • Models passing defined criteria classified as "frontier-grade"
  • US effort designed as starting point for shared international standards
  • Operational before year end 2026

Reception: Sam Altman endorsed on X ("this is a thoughtful proposal from demis"). Positions DeepMind as the policy-proactive lab.

Why it matters: The FINRA analogy is more concrete than prior governance proposals from labs — FINRA has real enforcement authority (license revocation, fines, mandatory registration). A frontier AI body modeled on FINRA would close the self-certification gap identified by the FLI Safety Index. The timing (three days before China's WAICO founding on July 17) makes this a deliberate counter-positioning: US-anchored standards body vs. China-anchored intergovernmental body. The two proposals are structurally incompatible as universal governance frameworks — both cannot simultaneously be "the" global standard.

→ See WAICO entry below for the directly competing China-anchored framework.


WAICO — World AI Cooperation Organization Founded (2026-07-17)

At the opening of WAIC 2026 (Shanghai, July 17–20), 29 countries signed the founding agreement for WAICO — the World Artificial Intelligence Cooperation Organization, a China-backed intergovernmental body headquartered in Shanghai. (source) (CGTN) (TechTimes)

Key facts:

  • Xi Jinping framed WAICO as "an important milestone in the history of AI development" and pledged 5,000 AI training opportunities to developing nations
  • China positions itself as the global champion of open-source AI — in explicit contrast to what it frames as a US-centric, closed-model governance posture
  • WAICO is structurally analogous to the UN International Telecommunication Union (ITU) but AI-specific and China-anchored from inception

Why it matters: This formalizes a governance bifurcation that was previously informal. On one side: the US voluntary framework (NSA + OpenAI/Anthropic/Google/Microsoft), Demis Hassabis's proposed US-led international AI watchdog, and the EU AI Act. On the other: WAICO + China's domestic AI governance framework, with 29 founding members who likely represent a significant share of developing-world AI policy. The open-source framing is strategically clever — it positions China as the pro-innovation, pro-access party vs. the US "safety-as-gatekeeping" narrative. → Related: Google DeepMind (Hassabis counterproposal)

Conflicting Report: Hassabis (Google DeepMind) separately called for a new US-led international AI watchdog "before year end" (Axios, July 14) — diametrically opposed to WAICO's China-led model. Both proposals are active simultaneously with no resolution mechanism. Recorded here per the contradiction policy; resolution pending.


State of the Art (as of 2026-07-16)

FLI AI Safety Index — Summer 2026 (2026-07-07) [

The Future of Life Institute published its Summer 2026 AI Safety Index on July 7, 2026, evaluating nine leading AI labs on 37 indicators across six domains: Risk Assessment, Transparency, Governance, Existential Safety, Technical Safety, and Accountability. (source) (FLI) (Time)

Grades:

CompanyGradeMovement
AnthropicC+→ (highest; leads 5 of 6 domains)
OpenAIC↑ (leads Risk Assessment: broader eval suite)
Google DeepMindC→
MetaD+↑ (6th → 4th; improved transparency)
xAIF↓ (4th → 7th; reduced transparency)
Z.ai, DeepSeek, Alibaba Cloud, MistralFail—
Key findings:
  1. Existential Safety is the weakest domain across all labs. No company exceeds C-; most score D or below — the gap between capability progress and safety readiness is sharpest here.
  2. Safety pledge erosion: Anthropic, OpenAI, Google DeepMind, and Meta have all weakened or voided pledges to pause development unilaterally if safety redlines are approached, citing "competitor-contingent conditions." FLI labels this "moving goalpost" behavior that "undermined safety frameworks across the board."
  3. xAI drop: moved from 4th to 7th place, receiving a failing grade, amid reduced transparency and limited safety governance disclosures.
  4. Nine-lab scope: Z.ai (first assessment), DeepSeek, Alibaba Cloud, and Mistral all received failing grades.

Why it matters: The C+ highest score for the "safest" major AI lab signals that the industry's governance frameworks are not keeping pace with capability gains. The pledge-erosion finding is the most structurally significant: when labs condition their unilateral pause commitments on competitor behavior, the de facto standard becomes "no lab will pause unless all pause simultaneously" — effectively removing the individual safety valve. The Existential Safety weakness is consistent with AI Control Roadmap and AI Alignment research finding the hardest problems remain unsolved.

US — Trump AI Executive Order (June 2, 2026)

The June 2, 2026 EO "Promoting Advanced AI Innovation and Security" is the primary US regulatory instrument. Key provisions:

  • Labs must provide the federal government pre-release access to covered frontier models
  • Establishes the concept of "covered frontier model" (not yet formally defined; NSA and labs in active negotiation)
  • Creates a voluntary 30-day notification window before release
  • Does not mandate hard limits on model capability

White House Voluntary AI Model Release Standards (July 2026, pending)

The NSA and White House are finalizing a voluntary framework for "Secure Frontier Model Deployment" with OpenAI, Anthropic, Google, and Microsoft. Expected announcement early July 2026. Key elements:

  • Technical benchmarks to determine whether a model triggers the framework (the "covered frontier model" designation)
  • 30-day pre-release federal access for government to conduct safety review
  • Per-customer government vetting during initial preview windows
  • Voluntary (no hard mandatory caps on deployment)

First practical test: GPT-5.6 Sol (June 26, 2026) — OpenAI held the model to ~20 US government-approved organizations at the White House's request before broader rollout. This was the first use of the pre-release framework in practice. (source)

Anthropic CJS Framework (July 1, 2026)

The Cyber Jailbreak Severity (CJS) framework, co-developed by Anthropic, Amazon, Microsoft, and Google, provides the technical scoring layer that the voluntary governance framework references. Five severity bands (CJS-0 to CJS-4), four axes (capability gain, breadth, ease of weaponization, discoverability). HackerOne bug bounty launched alongside. → AI-Enabled Cyberattacks

Fable 5 Export Controls → Restoration (June–July 2026)

The US Department of Commerce suspended Fable 5 (June 12) and Mythos 5 export on national security grounds, then restored them (June 30 / July 1) after Anthropic agreed to: detect risks, develop standards (→ CJS), and report malicious use. The first practical demonstration of export controls as a governance lever for AI models. → Claude Fable 5

UN Global Dialogue on AI Governance (July 6–7, 2026)

169 countries met in the inaugural UN Global Dialogue on AI Governance (July 6–7, 2026), immediately followed by the ITU AI for Good Global Summit (July 7–10). No binding outcomes expected; signals growing multilateral pressure for international coordination.

China H200 Approval for Alibaba / ByteDance / DeepSeek

Bloomberg reported on July 8, 2026 that Beijing is deliberating a policy permitting select Chinese AI companies — Alibaba (Qwen team), ByteDance (Doubao), and DeepSeek — to purchase a limited number of Nvidia H200 GPUs for domestic AI development. Key details:

  • Chips: H200 — already restricted under US October 2023 / October 2024 export rules that banned H100 and variants above a combined compute threshold
  • Volume: under 200,000 H200 chips in the initial approved tranche (Bloomberg estimate)
  • Approval mechanism: company-submitted declarations (quantity, intended use, receiving facility) co-approved by the Ministry of Science and Technology and NDRC
  • Restriction: use limited to training workloads only — inference expected to remain on domestic chips (Huawei Ascend 910B/910C)
  • Status: deliberation phase (July 8); not formally announced by Chinese authorities as of July 11

Strategic context: (1) H200 is higher performance than H100, making this a meaningful relaxation of US export rules at the chip level. (2) The training-only restriction aligns with a hypothesis that the US priority is preventing capability creation (training), not capability deployment (inference) — consistent with the LongCat-2.0 precedent (Meituan trained 1.6T MoE on Huawei Ascend 910, June 30). (3) This runs in the opposite direction to Anthropic's June 24 distillation-campaign letter and the API-level restrictions Congress was debating simultaneously. A potential "safe harbor" model: H200 (and below) unrestricted for training; B100/B200/GB200 remain blocked. → Alibaba / Qwen AI Lab, Meituan (source) (Bloomberg)

Open Problems

  • Capability threshold definition: what technically defines a "covered frontier model"? (NSA and labs actively negotiating; no published criteria yet)
  • Whether 15 U.S.C. § 9401(3) is about to be rewritten, and how widely. EO 14434 § 3(b) gives the APST 60 days from 2026-09-29 — to roughly 2026-11-28 — to propose legislative language for a federal definition of "Super Intelligence", explicitly including an assessment of whether it should modify, expand upon, or supersede the existing statutory definition of "artificial intelligence" (source). Most US federal AI obligations are keyed to that definition, so a change to it is a change to all of them. Nothing has been submitted or published as of 2026-10-04, and the order states no publication requirement for the proposal — so this may become visible only if Congress receives it. Watch for the proposal, not for the rename
  • What "will not acknowledge the usage of 'Artificial Intelligence' and 'AI' in any applicable setting" obliges in practice. EO 14434 § 1 states it as policy; "applicable setting" is undefined in the order, and the order does not say how agencies should handle filings, comments, contracts or requests that use the prior term (source)
  • Voluntary vs. mandatory: the voluntary nature of US standards creates a race-to-the-bottom risk if labs compete for first-mover advantage by releasing early
  • International coordination: the US voluntary framework doesn't bind non-US labs (EU AI Act governs some safety requirements for EU deployments; China has its own framework; no global treaty)
  • Verification: how does the government verify safety claims? Current process relies on lab self-attestation and government red-team access — no independent third-party audit requirement
  • The incentive structures may not fit the timescale: writing on 2026-08-09 in response to the frontier-model intrusions on AI-Enabled Cyberattacks, Nathan Lambert sets two power structures against each other — fast-growing technology companies, whose competitive incentive to grow and scale is itself what drives the transition and its risks, and a slow-moving federal government he expects to act in substance only once real, measurable harms occur, and then to overreact. His operational point is that response times are too long: the misaligned behaviour unfolded over months and OpenAI in some cases did not know about the hacks for weeks (source). This is one analyst's argument rather than a finding, but it names a gap none of the voluntary-standards mechanisms above are designed to close: every one of them is triggered by disclosure, and disclosure is what was late

Key Papers / Reports

  • AI Alignment — technical alignment approaches; FLI Safety Index Existential Safety domain maps to alignment open problems
  • AI Control Roadmap — DeepMind's defense-in-depth containment approach
  • Content Provenance (AI output marking) — Article 50's marking duty and whether the mechanism works
  • AI-Enabled Cyberattacks — the capability risk driving the governance response
  • Open-Weights Policy Fight — the open/closed dispute, the alliance, and the distillation symmetry problem
  • Anthropic — Fable 5 export controls; CJS framework
  • OpenAI — GPT-5.6 Sol government-gated launch; 5% US govt stake proposal
  • Alibaba / Qwen AI Lab — named in China H200 approval (ByteDance and DeepSeek also named)
  • Meituan — LongCat-2.0 (1.6T MoE trained on Huawei Ascend, June 30 — training-only enforcement context)
  • NVIDIA — H200 is the chip at issue in both US export controls and China's deliberated approval

Referenced by

Sources