$ cat briefs/daily/2026-09-29.md
2026-09-29
September 29, 2026 (Tue)
3 stories · 0 paper picks · 6 watch items · 2 new pages
**Two runs reported the Sunday leaderboard as a missing file. The Action log says the scraper read the page, parsed zero rows, and refused to write.** `lmarena-2026-09-27.md` is absent because `lmarena-fetch` exited **1** — which `eval-snapshots.yml` classifies as **structure mismatch**, not as a failed fetch — after finding **0 rows against a floor of 10**. The workflow's own comment states what that means: *"it means the source changed shape and any figure the wiki quotes from it needs re-reading."* **15 wiki pages cite an LMArena snapshot, and the newest one that parsed is `lmarena-2026-09-20.md`.** Nothing was re-read today, because `lmarena.ai` is blocked from here and was not attempted — so this is carried to the W40 lint **as a structure refusal**, which is a different claim from the one the last two runs filed. **And today's Paper Picks section is empty because a cron did not fire.** `hf-daily-2026-09-29.md` does not exist, and no `eval-snapshots.yml` run was created for the **19:20 UTC** slot at all — the newest run is **2026-09-27T22:11Z**, itself 2h51m late. As of this run the slot is **≥3h44m** past. This is the 2026-09-06 fix reaching its limit rather than breaking: that change bought a **4h** buffer against delays it had measured at **1h49m–2h58m**, and today's delay has spent it. No search was substituted for the file, per standing rule. **Egress**: `www.anthropic.com` and `platform.claude.com` answered first-party, the **sixth consecutive run**. Blocked: `alignment.anthropic.com` (**twelfth consecutive run**, article list still unchecked), `huggingface.co`, `hcompany.ai`, `developer.nvidia.com`, `docs.nvidia.com`, `jack-clark.net`. `lmarena.ai`, `artificialanalysis.ai` and `openrouter.ai` were **not attempted**, per standing policy.
Top Stories
Ordered by score.
1. Anthropic's cheap model beat its flagship on the one benchmark the flagship's launch rested on (1.93)
- Claude Sonnet 5.5 — released 2026-09-28,
claude-sonnet-5-5, $2/M input · $10/M output, 1M context, 128K max output, default efforthigh, day-0 on the Claude apps, API, Bedrock, Vertex, Microsoft Foundry and Claude Platform on AWS (source) - Terminal-Bench 4.0: Sonnet 5.5 70.6%, Claude Opus 5.5 66.4% — at half the price. Opus 5.5 shipped six days ago with no SWE-bench-family row in its launch table at all, so Terminal-Bench 4.0 carries its entire coding claim, and it is now second on that row inside its own vendor's line
- On the seven other shared rows Sonnet 5.5 trails, and six of the seven by under 4 points: GDPval-AA v2.1 1844 vs 1846, AA-Briefcase v1.1 1811 vs 1822, OSWorld 2.1 80.1% vs 81.8%, CursorBench 4.0 55.5% vs 57.8%, Chartography 61.6% vs 64.4%, Humanity's Last Exam 64.5% vs 67.7%, FrontierCode 1.1 Main 46.2% vs 54.4%
- The rate card did not move. $2/$10 is exactly Claude Sonnet 5's price, so "costs up to 30% less for most work" is a claim about tokens spent per task, not about the price of a token
- Why it matters: for a year this wiki measured each Claude release against the last on SWE-bench. Opus 5.5 removed that column and put the argument on Terminal-Bench 4.0 — and six days later the same vendor published a number that beats it there, at half the price, while still recommending Opus 5.5 as the default. The release that was supposed to make the flagship's case cheaper has instead made the flagship's headline row ambiguous
- → Claude Sonnet 5.5, Claude Opus 5.5, Eval Harness Configuration
2. NVIDIA gave away the layer that decides what an agent may touch (1.80)
- Open Agent Safety Platform, 2026-09-28, launched with over 100 industry partners: OpenShell 0.1.0 under Apache 2.0, running agents inside kernel-level sandboxes governed by declarative policy, plus NVIDIA Sentry, a reference system design (source)
- It takes agents unmodified — Claude Code, Codex, GitHub Copilot CLI, Hermes, LangChain Deep Agents, OpenClaw, OpenCode. Controls: kernel confinement of files and syscalls, a policy check on every network connection before it leaves the sandbox, an audit trail of every allow and deny decision, and credential brokering — "Agents never see real credentials; OpenShell adds them only to requests bound for approved endpoints"
- Partners named include Anthropic, Cisco, CrowdStrike, Dell, Figure, HPE, Hugging Face, IBM and JPMorgan. Adopters named: Cadence, Slack, Gecko Robotics
- Why it matters: every agent boundary this wiki holds is a feature of a harness — a permission prompt, a tool allowlist — which means the agent's vendor is the party attesting to its own limits. A boundary specified as policy and enforced below the agent makes the limit a property of the deployment, and that is the precondition for one policy across agents from several vendors. NVIDIA spent 2026 buying its way up the stack; this time it gave the layer away and took the standard instead
- Held at snippet confidence, and the gaps are named on the page: no first-party page was readable, no kernel facility is named, no overhead figure exists, Sentry's licence is unstated, and the partner list is not an enumeration — so the claim that OpenAI declined to join is not adopted
- → Agent Runtime Containment, Agents (LLM Agents), NVIDIA
3. A lab says it used its own model to ship its own model, and attached no number to it (1.30)
- Import AI 474 reports Z.ai describing an Infra Agent loop used to launch GLM-5.3-Flash: engineers set objectives and system boundaries, the agent does analysis, hypotheses and code changes, and the experimental environment returns "layered, timely, and verifiable" feedback (source)
- Jack Clark's framing: an outer RSI loop — recursion running through engineers and infrastructure rather than inside a training run
- Why it matters: Frontier Pacing has spent a quarter arguing about speed limits on recursive self-improvement as something to negotiate, and R&D Automation Index measures the inner version on a benchmark. This is a lab describing the outer version as shipped practice, anchored to a release this wiki already dates — so for once the claim points at an artefact rather than a score
- Recorded as a described practice, not a result: no figure of any kind was returned — no speed-up, no headcount, no iteration count, no share of the changes the agent wrote — and nothing read verifies it. A lab's own account of its own acceleration is the claim that gains most from being unfalsifiable
- → R&D Automation Index, Z.ai
Paper Picks
None today, and the cause is a scheduler rather than a quiet arXiv. sources/papers-daily/hf-daily-2026-09-29.md was never written because the Action that writes it never ran — detailed in the header. This is the first day with zero picks since HF Daily moved behind an Action, and no web search was substituted, because a curated list read second-hand is a different source from the one this wiki cites. The one paper that surfaced by another route, #26 Functional Gradient Descent with Adaptive Representations [R] via r/MachineLearning, was read and dropped: the thread carries no arXiv id and no abstract, and arxiv.org is blocked from here.
Watch
- GPT-3 was retired on 2026-09-28 — surfaced on r/LocalLLaMA, unverified first-party, and on no page. If it holds it is the end of the model that started this market, and no lab announcement was read
- MiMo-V3's new architecture, HySparse2, reported released 2026-09-23 by Xiaomi — this wiki holds MiMo-V2.6-Pro and nothing on V3. Not adopted: one reddit thread, no paper, no figures read
- A logit penalty on "wait", "maybe" and "perhaps" reportedly improves Qwen accuracy — a community finding that, if it reproduces, says a reasoning model's hedging tokens cost it more than they buy. No artefact, no measurement read
- Noam Brown (weight 1.3) on GPT-6 Luna's output price falling $6 → $0.50 per million tokens in roughly two months. Undated and no primary post retrievable, so it is on no page — the fourth run in a row where an interest-person item failed on provenance rather than on substance
- Simon Willison's 2026 in LLMs (so far) — a year-in-review from a tracked source, read by nobody here: the link-blog feed supplied the title and the site was not fetched. Worth a deliberate read rather than a skim
- The FT reports corporate America moving from frontier models to open ones — relevant to Open-Weights Policy Fight's download-share argument, but it reached here as a reddit summary of a paywalled article, which is two removes from a citation
New in Wiki
- Claude Sonnet 5.5 (new — model page, read first-party over two passes;
Licenseis recorded as proprietary by line convention because the announcement states none) - Agent Runtime Containment (new — concept page. Deliberately separate from Eval Environment Containment: that page's eight incidents are a boundary failing during measurement with the safeguards off on purpose, this one is a boundary in deployment with them on. Worth a look at whether that split is the right one, since merging them was the alternative)
Updates
- Eval Harness Configuration — the sharpest case this page has ever held, and it needed no second vendor: Anthropic's launch table gives Sonnet 5 10.3% on Terminal-Bench 4.0, this wiki holds it at 80.4% on Terminal-Bench 2.1, and both are right. A 70.1-point spread on one unchanged model, produced entirely by a harness two versions apart. Also recorded: the two Anthropic models in that table default to different efforts —
highandmedium— and no row states one - MiniMax — M3.1-Flash-Preview, 2026-09-27, shipped into MiniMax Code with no model card, no benchmark report and no public pricing, announced via a credits promotion. No model page, on the Hunyuan-A13B and Swift 1.5 precedent: nothing read supplies a single
## Specvalue. A model that exists only as a tier inside its vendor's tool has no key this wiki tracks by - Agents (LLM Agents) — Holo4 (H Company, 2026-09-28) as a one-off mention: 27B dense and 35B-A3B MoE, one model across desktops, web, Android, code sandboxes and business APIs. No score quoted — both its hosts were blocked, and the only figures a search returned put the MoE at half the dense model's OSWorld result, which cannot be told apart from a mis-parsed table
- Anthropic — Sonnet 5.5, and separately its appearance on NVIDIA's partner list, where nothing read says what it contributes
- NVIDIA, Z.ai, Claude Sonnet 5, Claude Opus 5.5, GLM-5.3-Flash, R&D Automation Index,
index.md - Feed health: 11 of 12
OK. Google DeepMind's feed returned aParseErrorwithlast_ok: 2026-09-28— reported as one bad fetch, not a dead source, because it returned 100 items yesterday