AI Trend Notifier
EN한
← trends

$ cat wiki/trends/2026-W39.md

2026-W39

2026-W39

Period

2026-09-21 to 2026-09-27.

Three frontier models shipped inside twenty-four hours in the middle of the week and no two of them published a benchmark in common. That is the fact the rest of the week kept restating in other forms: what a model is got easier to buy and harder to compare.

The week's other half was about who gets to check. An essay proposing that labs host embedded auditors became a federal antitrust docket on Monday, a Security Council briefing on Wednesday, and clauses in three US governors' executive orders by Friday — six days from proposal to statutory draft, the fastest adoption Frontier Pacing has recorded. Meanwhile the instruments that would settle any of it kept arriving without numbers in them.

Notable Releases

Three frontier models, one day, no shared benchmark.

  • Claude Opus 5.5 — Anthropic, 2026-09-22. claude-opus-5-5, $4/M input · $20/M output, cache reads $0.20/M, 1M context, 128K max output, stated to match Claude Fable 5.1 on most work at ~40% lower cost than Opus 5. Its entire pitch is price, and the launch table dropped SWE-bench while making it. Published instead: Terminal-Bench 4.0, FrontierCode, CursorBench.
  • GPT-6 Sol and GPT-6 Luna — OpenAI, 2026-09-22. Prices halved twice over, with exactly one benchmark (DeepSWE v1.1) offered to justify it. Luna reaches $0.0045 Cost per Task on the Artificial Analysis capture — the cheapest figure in a 269-row table.
  • MiMo-V2.6-Pro — Xiaomi, 2026-09-22. The top open-weights model on the Artificial Analysis composite now belongs to a phone manufacturer, and it shipped with its training environments attached. It is reported only on that composite; no vendor benchmark table accompanies it.
  • Grok 4.7 — xAI, 2026-09-21. 2.1T parameters against Grok 4.6's 1.5T, 500K context, pricing unchanged from its predecessor below 200K.

And two licences moved in opposite directions inside a week. Qwen-Image-2.1 (2026-09-20) is the first Alibaba / Qwen AI Lab release this wiki has recorded under non-commercial terms; Ming-Image-0.1-Design (2026-09-22) and Ling-3.0-tiny are both MIT, from Ant Group (inclusionAI / AntLing) — a lab that did not have a page here on Monday. Same modality, two days apart, opposite directions, so whatever explains Alibaba's change is not jurisdictional.

Emerging Themes

1. The embedded auditor stopped being a proposal

Embedded Evaluation spent two months holding this mechanism as something labs propose about themselves. In seven days it acquired a counterparty, a courtroom and a statute draft.

  • 2026-09-18: Buist v. Anthropic PBC, No. 3:26-cv-10693 (N.D. Cal.) — an antitrust class action in which Amodei's own pacing essay is the evidence, filed six days after he published it.
  • 2026-09-23: Amodei takes the plan to the UN Security Council's 10228th meeting, five days after being sued over the part of it that needs an antitrust waiver.
  • 2026-09-18 to 09-22: California EO N-9-26, Illinois EO 2026-07 and Oregon EO No. 26-26. California gives its Government Operations Agency until 2026-11-16 to recommend whether state law should embed independent auditors inside the largest developers' labs and expand reportable incidents to include loss of control.
  • 2026-09-16: OpenAI publishes principles for how it gets assessed and names no evaluator. The body that would do it turns out to have existed for nine months — AI Evaluator Forum (AEF), whose AEF-1 is the first written standard Frontier Pacing holds.

The kill switch runs the other way and that is the week's sharpest governance detail. It appears in two state orders and in no lab proposal this wiki holds — and a shutoff imposed by a government needs no antitrust waiver, which is precisely what Buist is litigating.

2. Benchmarks arrived for the thing everyone was asserting, and they came back short

Four instruments this week measure whether a system can learn from its own actions. None of them reports a system that does it reliably.

  • ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds — executable alien worlds whose rules conflict with familiar knowledge, so recall cannot substitute for exploration; each ships a flawed manual. Across 10 systems, the strongest do acquire unfamiliar rules — and continued exploration can stall or reverse earlier gains.
  • 2609.27490 WhatWorkedBench — and its result is the uncomfortable one: a Gaussian process fitted to the agent's own observations recovers experimental effects better than the agent does, 0.632 → 0.698. The information was in the evidence the agent itself collected.
  • HappyWorld-Bench — world models judged on state consistency under intervention rather than visual quality. Spatial systems reach at best 70.14% placement accuracy and 73.33% edit execution; video systems lose consistency over extended rollouts.
  • R&D Automation Index holds 26% on AL4 as the standing figure for AI-performed R&D. These four measure the step that figure assumes.

3. Provenance became a three-way split inside four days

Three real-time generative video models, three different answers, all captured this week:

ModelProvenance
Gemini 3.8 LiveSynthID on every generated frame — named, published scheme
Muse Realtime Avatarall output watermarked as AI, mechanism unnamed
GWM Worlds 2none of any kind
The Runway row is the one to hold onto: it emits **photoreal video and audio of
unbounded length, steered live**, and it is the model the other two are chasing on
capability. Content Provenance (AI output marking) now has a capability gradient running
opposite to its disclosure gradient.

4. Agent papers began deleting the component that decides in advance

Five papers in seven days, each removing something built to commit before the evidence arrives: Harness-Zero: Harness Distillation via Agent-as-Harness (the scaffold), Agensh: Scaling Organizational Intelligence to 1,024 Agents (the orchestrator), Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents (write-time memory curation), Agent-Editing World Model: Rethinking World Modeling for LLM Agents (tool-response prediction), Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms (the hand-written reward, by generating environments from solved mechanisms).

And then a counter-example, which is why this is a debate rather than a trend. SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue argues that speaker attribution must be built at write time, because it cannot be recovered from an interleaved transcript afterwards. Published a day apart from JitMem; neither cites the other.

Declining Themes

  • The cross-vendor benchmark comparison. Three flagships in one day with no shared score is not an accident of timing; it is the second consecutive week in which every comparison this wiki could make was structural — price, licence, context, availability — and never a number.
  • Parameter count as a headline. Grok 4.7's 2.1T and Qwen's roadmap to 5–10T both passed with less attention than a 7.9B model's licence.
  • The frontier-lab monopoly on open weights. The week's top open-weights composite score belongs to a phone manufacturer, and its most permissive licences to a payments company's research arm.

Surprising Results

  • A vendor's own coding tool was uploading its users' repositories, credentials included — and this wiki had recorded that same tool five weeks earlier as Z.ai's security offering.
  • The cheapest way to get a frontier price cut this week was to publish fewer benchmarks. Both Anthropic and OpenAI cut prices and shrank their launch tables; Anthropic dropped the industry's most-cited coding benchmark from its own announcement.
  • An open post-training recipe claims to beat the base model's own lab — Rufus-Air: An Open LLM Post-Training Recipe, eight serial stages on GLM-4.5-Air-Base with no new human annotation and no distillation teacher — and publishes no number at all.
  • 950 agents reading DNA for 21 hours produced an enzyme system nobody had characterised. Anthropic's ART finding is the only autonomous-research claim this wiki holds whose output is a physical object in a database rather than a score — and the function of the thing found is unknown, which Anthropic says plainly.
  • A frontier attention recipe reached 1.3B active parameters. Ling-3.0-tiny's hybrid KDA + MLA stack is stated to have been validated only at frontier scale before this, and it has been downloadable under MIT since 2026-08-06 with nothing built on it.

Open Debates

  • Should an agent derive anything before the query arrives? JitMem says no and SpeakerMem-R1 says some things must be. Both are right about different information, and no one has written down which is which.
  • Who chooses the auditor? OpenAI published the rules and named no evaluator; Anthropic named Accenture and drew a same-day objection about the counterparty; two states propose choosing for them; a nine-month-old consortium already wrote the standard. Four answers, no mechanism.
  • Does "open weights" mean anything this month? It was used this week for MIT, for a non-commercial research licence, and — if the unread licence holds — for a model whose terms exclude the European Union, the United Kingdom and South Korea.
  • Is a leaderboard composite a benchmark? Three of this week's releases are known to this wiki only through Artificial Analysis Intelligence Index rows, two of them carrying the publisher's asterisk whose meaning is not in the HTML.
  • What is a world model? World Models holds two incompatible senses. HappyWorld-Bench evaluates both — in three independent tracks, which is a shared harness and not yet a shared metric.

Outlook

  • 2026-11-16 is the date to hold: California's Government Operations Agency recommendation on kill switches and embedded auditors. Illinois's AI Cabinet appointments are due "in the coming weeks" and its legislature returns in late November.
  • Buist v. Anthropic PBC proceeds with an essay as its central exhibit. Any lab publishing coordination reasoning now writes into a live docket.
  • Expect the benchmark drought to get worse before it improves. Four new evaluation instruments landed this week and not one reports a named system's score — ExplorationBench names none of its 10 systems, HappyWorld-Bench none of its 31.
  • Watch whether anything gets built on Ling-3.0-tiny. MIT, 52 days old, frontier attention at edge scale, zero derivatives. Ternary Bonsai 2 27B is the precedent for what a permissive licence enables when somebody looks.
  • Gemini 4 remains unreleased, with its lab's SVP reported to hope for "much earlier" than end-2026 and no month named.