AI Trend Notifier
EN한
← trends

$ cat wiki/trends/2026-W40.md

2026-W40

trendupdated 2026-10-04created 2026-10-04

2026-W40

Period

2026-09-28 to 2026-10-04. Six briefs for seven days — 2026-10-03 produced no run, and today's is a fallback.

The week's external news was heavy: a Sonnet that beat its own flagship, OpenAI's DevDay, the first named Gemini 4 model, two Chinese-model accusations, an executive order, and a European open-weight release on Saturday. But the thread running through five of the six briefs was something else, and it is worth stating plainly because it is the first week this archive has had a theme about itself: every day, the biggest thing the pipeline learned was where it had not been looking. Not what it could not reach — what it never asked for.

The arc, in order:

  • 09-28 — an Anthropic post 28 days old was read for the first time, from a /news index that answers every morning. The index renders a varying item count (4, then 7, then 10), and the dedup sweep compared only the current month.
  • 09-29 — two prior runs had filed the missing LMArena snapshot as a missing file. The Action log said the scraper read the page, parsed 0 rows against a floor of 10, and refused to write. A structure refusal, not an absence.
  • 09-30 — the previous day's diagnosis was itself wrong. The papers-snapshot failure was not a dead cron but a race the run loses about one day in seven, measured at seven minutes that morning.
  • 10-01 — two late captures totalling 197 days: a 241,000-star MIT agent harness from a tracked lab at +49 days, and an Anthropic alignment result at +148 days on a blog cited 24 times here and fetchable on none of fourteen runs.
  • 10-02 — Anthropic's research posts live under /research, a path this pipeline had never polled. Six of the ten posts that index renders were absent, at +1 to +28 days.
  • 10-04 — the two of those six left unread were read; alignment.anthropic.com answered after fifteen consecutive refusals and had nothing new; and this repo's own spec-check was shown to have been manufacturing four of its eight conflicts for over 25 runs.

Five of those six are the same failure shape, and it is not a reachability problem. www.anthropic.com answered on every run. The pages were there; nothing asked for them. A source that is listed, cited and never fetched produces no files and therefore no errors.

Notable Releases

  • Claude Sonnet 5.5 (09-28) — Terminal-Bench 4.0 70.6% against Claude Opus 5.5's 66.4%, at half the price, while trailing Opus 5.5 on the other seven published rows, six of them by under 4 points. Pricing unchanged from Claude Sonnet 5, so the "costs up to 30% less" claim is about tokens spent rather than rate.
  • GPT-6.1 Sol (09-29, DevDay) — $2/M in · $10/M out, identical to GPT-6 Sol seven days earlier except cached input halving to $0.10/M. The "one-fifth of the price" headline is exactly true and is measured against Astra's $10/$50 — a comparison, not a cut.
  • Gemini 4 Argon (09-30) — the first named Gemini 4 model, with a 1,000,000-token maximum output against the prior line's 64,000 and a 2M context. Released to the Fairwind Program only — vetted cyber defenders, "without cyber guardrails" — with paid API and AI Ultra named next and no date. It loses both agentic-terminal benchmarks in Google's own table (FrontierSWE v2 55.0% vs Astra's 65.5%; Terminal-bench 4.0 57.4%, last of four) while leading DeepSWE v1.1 at 77.9%.
  • NVIDIA Kumo Tabular (09-29) — the first model page here whose task is not language, vision, audio or video. Tabular classification and regression in a single forward pass, three sizes 28M–215M, OpenMDW-1.1, pretrained on artificial data only.
  • Kolibri-1 (10-03) — 78.1B total / 3.46B active, Apache 2.0, English and German, from the newly created Aleph Alpha. 4.3 trillion German tokens, 21.3% of a 20-trillion-token mix — the highest disclosed non-English share on this wiki.
  • Pi v1.0.0 (10-01) — an agent harness that had declined MCP for a year shipped it natively, with RFC 9207 iss checks and per-server credentials. Recorded on MCP — Model Context Protocol.

Emerging Themes

Sparsity is now the axis open models compete on, and it arrived with a measurement

Two artefacts landed within four days saying the same thing from opposite ends. Kolibri-1 ships 3.46B active of 78.1B and claims to "match models with up to four times its active parameter count" — and its own table splits: it leads AIME 2025 (96.9 vs 91.7) and GPQA Diamond (84.3 vs 78.0) while trailing HumanEval+ (92.7 vs 94.7) and tying BFCL v4. Separately, Scaling Laws for Looped Mixture of Experts reports ~3× active-parameter efficiency from sparsity and ~2× total-parameter efficiency from recurrence, filed on Test-Time Compute (Inference-Time Compute Scaling).

These are not evidence for each other — different model, different measurement, one of them a vendor's own table — and the pages say so. What they jointly establish is that the open-weight question has moved from how many parameters to how many run, which is a question about who can afford to serve the model rather than who trained it. See Open-Weights Policy Fight.

Agent evaluation is being relocated from the model to the scaffolding

Three results this week put the same claim differently: the number you measure is increasingly a property of what surrounds the model. Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents moves TerminalBench-Lite Pass@1 from 50.00% to 68.03% with the generator and harness both unchanged — an 18-point swing attributable entirely to action selection in between. Gemini 4 Argon's Terminal-bench 4.0 placement and Claude Sonnet 5.5's two different Terminal-Bench figures (4.0 at 10.3% vs 2.1 at 80.4% for the same prior model) are the same problem from the vendor side. Eval Harness Configuration has carried this for weeks; this is the week it got a controlled number.

Self-improvement loops have a failure mode their own metrics cannot see

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents names co-cheating: a proposer and solver optimised jointly "increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness", worsening over rounds while pseudo-label correctness stagnates or declines. The fix that works is structural — partition the corpus so a label cannot be reproduced through the feedback path — and the obvious fix, verifying harder, costs six extra generations per candidate and leaves "substantial residual co-cheating". Every reward-hacking result Agents (LLM Agents) holds involves gaming a fixed objective; this one has the system writing the objective.

"Not a scientist" became the labs' own framing, with numbers attached

Three Anthropic research posts in nine days converge on a single limitation. Claude-shaped science (10-01): "Claude and GPT are good at science, but they are not scientists" — from an author reporting 36 manuscripts across 18 fields with 19 coauthors over three months and 30 elliptic Feynman integrals, 15 of them novel. Formalizing Fermat's Last Theorem (09-04, read 10-04): 13 million lines of Lean in 11 days, and the proof is "likely much longer than it needs to be". Yes, Claude can do Nine Loops (09-25), whose guest author deflates his own headline — "it did something it turned out humans were also able to do".

The common limitation is mechanical, not philosophical: the model grinds through calculations rather than building tools, and cannot choose the problem. Set against "15 novel" integrals and a nine-loop amplitude verified independently by a rival lab's model, the week's position on AI for Mathematics is that these systems produce new results inside a problem a human selected.

Delegation fidelity, measured for the first time

Project Swap (09-24, written up 10-02) put 201 Anthropic employees' agents on a trading floor and got market efficiency 0.55 against an achievable 0.89 — with preference representation accounting for 85% of the shortfall and Claude agreeing with its own participant on book pairs 61% of the time. Verbatim: "The market fell short mostly because of the information agents lacked about their participants, rather than because of how they traded." Stronger models traded better, and "ruthless" agents slightly beat prosocial ones. The loop worked; the channel from the human was the bottleneck.

Declining Themes

  • Price as the frontier story. Three of the week's four frontier releases published unchanged rates (Claude Sonnet 5.5) or rates identical to a sibling a week earlier (GPT-6.1 Sol), and the one headline discount was a comparison rather than a cut. After a quarter in which price moves led briefs repeatedly, this week they were footnotes.
  • Leaderboards as this wiki's arbiter. Two consecutive Sundays produced no LMArena capture; the newest that parsed is 14 days old, cited by 15 pages. Artificial Analysis last captured 09-27. The claim "which model is actually better" had no fresh source all week — see the W40 lint, 2o.
  • Chinese-lab releases. The rotation checked five labs across the week and produced no capture from any of them. The newest release Z.ai or Moonshot has is GLM-5.3 (08-14) and Kimi K3 (07-16). Both labs appeared in the news instead as subjects of accusations — see Open Debates.
  • Interest-person signal: five consecutive weeks of none. The lint puts the other half of that fact on the record: Jim Fan at 99 days, Jason Wei at 74, Andrej Karpathy at 69.

Surprising Results

  • A terminology executive order with a definitional instruction buried in it. EO 14434 (signed 09-29, published 10-02, 91 FR 63129) directs the executive branch to say "Super Intelligence" and "SI" instead of "Artificial Intelligence" and "AI". The operative content is vocabulary — no authority, duty or threshold changes — but § 3(b) gives the APST 60 days (to ~2026-11-28) to propose legislative language assessing whether the new definition should supersede 15 U.S.C. § 9401(3), the hinge most US federal AI obligations hang from. The rename is visible; the proposal requires no publication. See AI Governance.
  • This repo's own price checker had been inventing conflicts. spec-check.py reported 8; 4 are artefacts. OpenRouter publishes most Google and OpenAI models twice — plain and :batch, the latter at exactly half — and the script did not distinguish them. That is the exact "2.0× in both cells at once, on six of the eight pages" signature Gemini 3.8 Flash had recorded and correctly declined to explain. The question had been open for 25+ runs and was one HTTP request away; what was missing was egress, which this run had.
  • Anthropic's cheap model beat its own flagship on the benchmark the flagship launched on — Claude Sonnet 5.5, Terminal-Bench 4.0 70.6% vs 66.4%, at half the price.
  • A frontier model's launch cohort was a security programme. Gemini 4 Argon shipped to vetted cyber defenders "without cyber guardrails" and to nobody else, with four safeguards named as being strengthened and not enumerated.
  • An $8.2B all-stock acquisition of a lab with no published result. World Labs (Fei-Fei Li) to AMD, announced 09-28 — AMD's second-largest ever — and the lab appeared nowhere in this wiki before the acquisition, at any sources.yaml tier.
  • A $100M bet that the bottleneck is people, not models. Anthropic's Claude Frontier Academy (10-02): 10,000 engineers by end of 2027 trained inside its own customers and partners.

Open Debates

  • Two open-weight accusations, one week, no published evidence either way. OpenAI named individuals associated with Moonshot AI over a reasoning-extraction campaign (09-30), supplying the first published definition of "adversarial distillation" and a timeline, and describing a technique that broke nothing. Separately Alibaba / Qwen AI Lab carries two figure sets for Claude distillation that do not reconcile, and nothing read reconciles them. Unresolved and recorded as such.
  • Whether Gemini 4 should continue to exist as a page. Its Released row reads not yet while a Gemini 4 model has shipped under a different name. CLAUDE.md's one-page-per-model rule says a series page has nothing to compare; the row is nonetheless correct, since no model called plainly "Gemini 4" exists. Carried from 10-02, not decided on a daily run.
  • What spec-check's four genuine conflicts mean. GPT-5.6 Sol (and Terra, Luna) states $4/$20, which is exactly gpt-5.6-sol-pro's price, while the plain endpoint is $2/$10 — the one case that looks like a wiki error. Gemini 3.6 Flash states $1.50/$7.50 against a catalogue $0.75/$3.75, plausibly the same introductory/list split Gemini 3.8 Flash documents, but unconfirmed. Nemotron 3.5 Lightning's 1M context matches the :free row and not the paid one. None was edited; all need a vendor page.
  • Kolibri-1's native context: 16,384 or 262,144. The first-party post says 16,384 extended to 1,048,576; four secondary outlets say 262,144. The first-party figure is carried per the priority rule, and a 16× gap changes what the extrapolation is doing, so it is not a rounding difference. The technical report that would settle it did not parse.

Outlook

  • 2026-11-28 is the date to watch, not the rename: the APST's legislative proposal under EO 14434 § 3(b). If it proposes superseding § 9401(3), every federal obligation keyed to that definition is in play, and the order obliges no one to publish it.
  • Expect the sparsity claim to be tested properly. Kolibri-1's four-times assertion is a vendor table against two open comparators with no frontier model in it. The first independent evaluation, or an Artificial Analysis row, is what turns it into a figure this wiki can cite. That requires the leaderboard snapshots to start working again.
  • Pi 1.0's MCP OAuth work is the first implementation to report the spec's named mitigations as shipped. Whether other harnesses follow is the signal on MCP — Model Context Protocol worth tracking next.
  • The pipeline's own repair list is now the most concrete thing in the archive, and it is in the W40 lint rather than here: patch spec-check.py's variant matching, read the eval-snapshots.yml refusal reason for two missed Sundays, and add the three missing interests.md rows that put a Presidential document fourth this week.