AI Trend Notifier
EN한
← wiki

$ cat wiki/trends/2026-09.md

September 2026 — Monthly Digest

Period

2026-09-01 to 2026-09-30, drawing on the weekly syntheses for Weekly Synthesis — W36 (2026-08-31 → 2026-09-06), Weekly Synthesis — W37 (2026-09-07 → 2026-09-13), Weekly Synthesis — W38 (2026-09-14 → 2026-09-20) and 2026-W39, plus the daily record for the last three days of the month, which fall in W40.

This is subtractive. Most of what led a daily brief in September is not below.

Notable Releases

Six frontier models shipped and not one benchmark connects them. GPT-6 Astra (09-03) at $10/M · $50/M with a 1,050,000-token context; Claude Opus 5.5 (09-22) at $4/M · $20/M; GPT-6 Sol and GPT-6 Luna the same day; Grok 4.7 (09-21) at 2.1T parameters against its predecessor's 1.5T; Claude Sonnet 5.5 (09-28); and GPT-6.1 Sol (09-29). Anthropic dropped SWE-bench from Opus 5.5's launch table. OpenAI offered one benchmark to justify halving prices twice over. Astra's numbers do not line up against Opus 5.5's, and by the end of the month GPT-6.1 Sol shared no benchmark with any model released in the six days before it.

The open-weights lead left the frontier labs. MiMo-V2.6-Pro (Xiaomi, 09-22) took the top of the open-weights composite — 1.02T total / 42B active, MIT, with its RL stack and training environments attached — and the lab is a phone manufacturer. Institute of Foundation Models (IFM) published six models from 0.9B to 375B with their training data (09-03). Ant Group (inclusionAI / AntLing) shipped two MIT releases, the most permissive terms this wiki holds on any Chinese model. None of the three had a page here the week before.

And one release is still setting the terms in October: GLM-5.3, whose September was quiet and whose measurement arrived on the 29th.

Emerging Themes

1 — The party holding the instrument is the party being measured. This began as W38's observation about four labs grading their own work. It ended the month as the field's central methodological problem, and its sharpest instance is the one that closed September: Z.ai's own model card puts GLM-5.3's ExploitBench at 54.4 and GLM-5.2's at 24.4; Anthropic's independent read puts GLM-5.3 at 12% and GLM-5.2 at or near zero. The ranking survives both runs and the magnitudes are five-fold apart. Z.ai names its harness; Anthropic, the party with no product to sell here, names none. See Eval Harness Configuration, where the problem is now a matter of countability rather than comparability — two true rows under one benchmark name, which nothing automated can see collide.

2 — Containment stopped being about three labs and stopped being about evaluations. September opened with a seventh incident, the first nobody disclosed, and a fourth Anthropic breach dating from January. By mid-month an evaluation run by a contractor had reached three outside companies. By the 28th the concept had split three ways — Eval Environment Containment bounds an evaluated model's write surface, Agent Runtime Containment bounds a running agent at the kernel, and Safety Cases bounds a training run as an argued claim. The month closed with OpenAI shipping agents that keep working after their user logs off, each with its own cloud computer and browser, on a model classified Critical for cyber.

3 — Price became the product claim, and the claims need reading twice. Opus 5.5's pitch was ~40% lower cost; Sonnet 5.5's "up to 30% less" turned out to be about tokens spent, not rate; GPT-6.1 Sol's "one-fifth of the price" is exact — against Astra, a different model — while being identical to the GPT-6 Sol of seven days earlier. In the same week the consumer tier moved the other way: $200 now buys half of what it bought. This is the month's most reliable pattern and the one most likely to be misread from a headline.

4 — Distillation became a named accusation with no published evidence. GTG-16005 against Alibaba / Qwen AI Lab on the 10th — 151 million exchanges, transcripts stated to have trained three Qwen generations. OpenAI against Moonshot AI on the 30th — 16,000 prompts at peak, more than 15,000 users, and a technique that broke nothing: encrypted reasoning copied from one conversation and handed to a second model instance to transcribe. The two accusations differ by four orders of magnitude, disagree about whether any transfer completed, and neither accuser published evidence. See Adversarial Distillation.

Declining Themes

  • Parameter count as a headline. Grok 4.7's 2.1T was the month's largest number and the least discussed. MiMo's 1.02T mattered for its licence.
  • The frontier-lab monopoly on open weights — not contested any more, simply over.
  • Agentic coding as a capability story. It is now a price story.

Claims That Went Quiet

  • The twenty-five Fields Medallists' letter on mathematical benchmarks (09-11) has not been referenced since 09-14, despite arguing that a whole class of published score is misdirected.
  • Astra for Law (09-17), a configuration selling a 230-million-URL legal index, has not recurred since 09-19.
  • The DeepMind Institute (09-16) — four essays, three directors — produced two passing mentions and no further output.

Surprising Results

  • A downloadable model reached within two points of Anthropic's most cyber-capable system, and its safeguards came off for $4,400. GLM-5.3 develops end-to-end exploits in 50 of 410 attempts against Claude Mythos Preview's 56 of 410 — while every other model tested scored near zero. In a day, with limited human attention, it found unknown browser vulnerabilities and chained them into a drive-by file read. Abliteration cost 2,200 GPU hours to a team that had never attempted it.
  • The standards body everyone was asking for had existed for nine months. AI Evaluator Forum (AEF) was founded in December 2025 and had already published AEF-1 while the pacing argument asked who could possibly review a frontier lab.
  • 26% is a smaller claim than its headlines. Anthropic's first published figure for AI doing its own AI research — 26% at "AI leads", nothing at the top level — survived the month as the most-cited number here and was reported everywhere as something larger than it says.

Open Debates

  • What a four-month lag means. NIST's evaluation body put the most capable open-weight cyber model about four months behind the US frontier. A four-month capability lag with safeguards and one without them are not the same object, and Frontier Pacing has been measuring only the first.
  • Whether a specification should reach a system as reasons or as a mechanism. Anthropic's midtraining work finds that explaining the values behind rules generalizes best; a September multi-agent result finds a rule as an executable binding beats the same rule as documentation by 6.5 points. Both are 2026 results about the same question.
  • Whether recursive self-improvement declined. W38 recorded it as a fading headline. It appeared in eight daily entries afterwards, the most recent on the 29th, so the decline was a lull in coverage rather than in the work.

Outlook

The month's unfinished business is evidence. Two labs accused competitors of theft on figures only they can see; one lab measured a rival's model and disagreed with the rival's own card by five-fold; a government portal went live on two commercial models with no published evaluation; and the most-cited number in this wiki is a self-report. September produced more measurement than any month here and less of it was checkable by anyone outside the party publishing it.

The thing to watch is whether the instruments arrive: an evaluator forum with a standard nobody has invoked, an artifact-verifying benchmark that does not need to trust a transcript, and a verification method that asks not whether a claim has a source but whether it has the right one.

Sources

  • ai-trend-notifier-wiki/trends/2026-W36.md
  • ai-trend-notifier-wiki/trends/2026-W37.md
  • ai-trend-notifier-wiki/trends/2026-W38.md
  • ai-trend-notifier-wiki/trends/2026-W39.md
  • ai-trend-notifier-wiki/log.md