$ cat briefs/daily/2026-09-30.md
2026-09-30
September 30, 2026 (Wed)
5 stories · 3 paper picks · 4 watch items · 8 new pages
**Yesterday's brief blamed a cron that had not failed, and the real defect is a seven-minute margin.** Yesterday this section reported that no `eval-snapshots.yml` run existed for the papers snapshot's slot and called it the 2026-09-06 buffer running out. Both halves were wrong. The slot is **07:20 KST = 22:20 UTC**, not the 19:20 UTC that measurement was taken against, and **run 63 started 2026-09-28T23:49:19Z** — **45 minutes after** yesterday's run read the directory. The last seven scheduled runs started at **22:17 · 22:24 · 22:20 · 21:53 · 22:11 · 23:49 · 22:57 UTC**: ordinary GitHub queueing of 0–90 minutes, with only one landing behind the daily run. Today's file was committed at **22:57:44Z** and this run began at **23:04 UTC**. **Seven minutes.** The papers intake is a race the run loses roughly one day in seven, and when it loses, the brief ships with no Paper Picks — as it did yesterday. That is a scheduling margin with a one-line fix, not a broken cron, and it is carried to the W40 lint as such. **The Sunday leaderboard hole is unchanged and is now ten days wide.** `lmarena-2026-09-27.md` is still absent for the reason established yesterday — the scraper parsed **0 rows against a floor of 10** and refused — and the newest capture that parsed is still **`lmarena-2026-09-20.md`**, cited by **15 wiki pages**. Nothing was re-read: `lmarena.ai` is blocked from here and was not attempted. **Egress**: `www.anthropic.com` answered first-party for the **seventh consecutive run** and carried nothing new — its news index renders 10 items, newest **2026-09-23**. Blocked: **`openai.com`** (which is why today's three biggest stories rest on search rather than on a vendor page), `simonwillison.net`, `huggingface.co`, and `alignment.anthropic.com` for the **thirteenth consecutive run**. `lmarena.ai`, `artificialanalysis.ai` and `openrouter.ai` **not attempted**, per standing policy.
Top Stories
Ordered by score.
1. OpenAI proposes that a frontier training run should need paperwork to continue — and does not say it obeys it (1.99)
- Towards safety cases for frontier AI training, 2026-09-28. The sentence is the whole of it: "structured safety documentation should be required before continuing any frontier reinforcement learning training run" (source)
- The instrument is borrowed from aviation and nuclear power, where a regulator must accept the argument before operation. A case should cover three layers — alignment training, containment, monitoring
- Five process recommendations: objection rehearsals, executive veto power, accountable training leads, default-to-shutdown on failure, data rollback. Four of the five are about who may stop a run — only data rollback is an engineering capability
- What is missing is the part that decides whether it matters. "Should be required" is the grammar of a proposal. Nothing read says OpenAI requires this of itself, names a run it was applied to, names a reviewer, or names who holds the veto. The nearest candidate reviewer — AI Evaluator Forum (AEF)'s AEF-1, which this wiki established had existed for nine months while everyone in this argument asked for a body — is not mentioned. One search pass, no first-party read
- Why it matters: every gate this wiki tracks, Preparedness Framework included, gates a deployment on an evaluation result. This gates continuing to train on a document, which is a far earlier intervention point and the one Frontier Pacing has been arguing about since July without anyone naming a mechanism. It is also Altman's Security Council sentence of five days earlier — "should not train models unless they can make a strong case" — with a form attached
- → Safety Cases (new), Frontier Pacing, Agent Runtime Containment
2. OpenAI shipped an agent with no session boundary, and published no evaluation of it whatsoever (1.85)
- dots, announced 2026-09-29 at DevDay 2026: always-on agents inside ChatGPT. Each one runs on GPT-6 Astra, is given its own cloud computer and browser, and is stated to keep working after its user logs off (source)
- Reachable in the ChatGPT desktop app, mobile app and web, in Slack or Microsoft Teams, and by voice call; plugins stated to connect to more than 4,000 apps. Absent from Free and the $8 Go plan
- The entire capability claim is "remarkably capable". No benchmark, success rate, evaluation, sample or failure analysis was published or returned by either search pass — for a product running unattended on a persistent browser, on the first model OpenAI confirmed at its Preparedness Framework's Critical cybersecurity threshold
- Why it matters: every agent product this wiki holds — Codex, ChatGPT Work, Claude Code, Grok Build — runs a task and stops. Removing the session boundary moves the interesting questions to what happens while nobody is watching: what it may spend, what it may send, what it may install. It arrives one day after NVIDIA shipped a kernel-level runtime that bounds exactly that, with 100+ partners, and OpenAI is not among the reported ones
- → Agents (LLM Agents), Astra, OpenAI
3. GPT-6.1 Sol's "one-fifth the price" is exactly true, and it is not a price cut (1.59)
- GPT-6.1 Sol — 2026-09-29,
gpt-6-1-sol, $2/M input · $10/M output · cached input $0.10/M, 1.05M context, 128K max output (source) - The arithmetic checks out against this repo. This wiki holds Astra at $10/M input · $50/M output, read first-party on 2026-09-03; $2/$10 is exactly one fifth of $10/$50. That is the first time a blocked-host capture here has been confirmed rather than merely corroborated, and it works only because the earlier read was kept
- It is also unchanged from GPT-6 Sol, released seven days earlier at $2/$10. The only rate that moved is cached input, $0.20 → $0.10. The headline compares to a higher tier, not to the model it upgrades
- Published benchmarks: OSWorld 2.0 71.4% against Astra's 73.5%; DeepSWE v1.1 parity claimed with no number for either model. No benchmark is shared with any model released in the six days before it. Same-day plan changes cut the other way: Pro 200's Codex and Work allowance drops from 20× to 10× the Plus level and GPT-6 Pro chat caps fall 200 → 100/week, while a new Pro 500 appears at $500/month
- Why it matters: this is the second lab in two days to headline a cost reduction that is not a reduction in the rate card — Claude Sonnet 5.5 did it on 2026-09-28, at the identical $2/$10. Two mid-tier models, two labs, one price, and neither announcement's cost claim survives being read as a rate. Meanwhile the same company halved what $200 buys on the consumer side
- → GPT-6.1 Sol (new), GPT-6 Sol, Astra
4. NVIDIA published a foundation model for spreadsheets, and gave it away (1.40)
- NVIDIA Kumo Tabular — 2026-09-29, three sizes 28M–215M, OpenMDW-1.1, pretrained on artificial data only. Predicts labels for unlabeled table rows in a single forward pass: no training, no tuning, no feature engineering (source)
- Reported 1st on TabArena (ELO 1950), BeyondArena (ELO 1418 / Improvability 7.78%), TALENT and ScoringBench, at 26× LimiX-2's speed on a single RTX 6000 Pro — all four placements the vendor's own reading, with no third-party reproduction
- Which of the three sizes posts those numbers is not stated, across an 8× parameter span. That is the Ling-3.0-tiny defect, where a sibling's score had been quoted as the model's. One search pass,
huggingface.coblocked - Why it matters: two NVIDIA releases in two days, both given away — Apache-2.0 OpenShell on 09-28, OpenMDW-1.1 weights today — against a 2026 record of buying and licensing its way up the stack. And this one reaches a workload class no model in this wiki has touched, where the incumbent is gradient-boosted trees rather than any lab. A note on weights:
interests.mdhas no topic row covering tabular prediction, so this was scored on an unmatched topic at 1.0; worth a line in the weights file either way - → NVIDIA Kumo Tabular (new), NVIDIA
5. An $8.2B valuation of world models, attached to no published result — from a lab this wiki had never recorded (1.30)
- World Labs — AMD to acquire it for $8.2 billion all-stock, announced 2026-09-28, close expected end of 2026 subject to regulatory approvals. AMD's second-largest acquisition on record, after roughly $50 billion for Xilinx in 2022. Founder Fei-Fei Li to become AMD EVP and chief scientist (source)
- Spatial-intelligence models that generate, reconstruct and simulate interactive 3D environments from text, image and video, plus robotic learning and simulation. Nine outlets agree on the terms and none names a model, licence, benchmark or headcount
- Why it matters: World Models has tracked this field since May and World Labs appears in none of its entries, nor in
sources.yamlat any tier. The field's visible artefacts this quarter are a research preview with no benchmark (GWM Worlds 2, announced not released) and a benchmark showing systems failing object permanence — so the price and the measured capability are being set by different processes. Captured at +2 days only because an AI newsletter covers semiconductor deals; Runway was +23 days four days ago for the same reason, and the gap is in which organisations get watched - → World Labs (new), World Models, AMD
Paper Picks
From hf-daily-2026-09-30.md — present today, unlike yesterday. Upvote counts are that community's popularity signal and nothing more.
Post-Training Leaves Behavioral Shadows on Unrelated Decisions — arXiv 2609.29233
- TL;DR: a capability can be transferred using one word of teacher output per prompt — no task examples, no logits, no teacher parameters. Pick prompts where teacher and student's shared public ancestor is nearly indifferent between two ordinary words, so the teacher's choice is attributable to its post-training rather than to the common prior. +5.34 pp on HumanEval+ (Qwen2.5-1.5B) over an exact nuisance-matched control, reproduced in science, commonsense and reading comprehension across generations, sizes and families.
- Why read it: every anti-distillation measure this wiki holds — Claude Opus 5.5's preserved thinking included — protects the content of outputs. This needs one word on prompts the extractor chose, which is indistinguishable from ordinary API traffic, and nothing read proposes a defence.
- → Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Relic: From Multi-Agent Collaboration to Persistent Organizational Capability — arXiv 2609.32965
- TL;DR: converts recurring multi-agent failures into organization-owned executable protocols bound to the runtime. The result that transfers is the comparison of the same rule as readable text against as an executable binding under fresh-member transfer: 25.4% none → 34.6% text → 41.2% bindings. Complete-contract delivery 14.06% → 19.76% over 360 runs, ten workloads, three models.
- Why read it: it measures the thing agent frameworks assert. A +6.5-point gap between documentation and enforcement is an argument about where coordination knowledge has to live, and it lands four days after OpenShell reached the same conclusion about safety policy from the opposite motive.
- → Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
Imprint Reader: From Weight-Update Readout to Behavioral Intervention — arXiv 2609.35261
- TL;DR: trains a model to describe a frozen weight update in natural language, with no-change and random-perturbation controls. The readout is weak and reported as weak — Pass@100 2% for knowledge, 16% for behavior — but the Reader is differentiable with respect to the update, and its gradients drive MetaEdit: harmful-prompt refusal 57.9% → 64.1% at a 0.5% pruning rate, BFCL Overall 41.69% → 44.60%, from behavior descriptions with no target-task training data.
- Why read it: the asymmetry is the finding — a signal too unreliable to trust as an explanation is still useful as a gradient. And read with the pick above, two independent methods say the same thing: a training update is legible from outside the run, whether you hold the weights or only an API key.
- → Imprint Reader: From Weight-Update Readout to Behavioral Intervention
Watch
Signals recorded without a page.
- Beijing is reported to have expanded travel restrictions on AI talent (Trivium China, 2026-09-29). Single source, reported rather than published, no instrument named — but a state limiting the movement of researchers is a different lever from the export controls this wiki tracks, and worth the next mention it gets
- The US and China are reported to have agreed to discuss AI risk (Trivium China, 2026-09-30), alongside a separate item on China "delicately rejecting" the word "superintelligence". No venue, date, scope or participant list in anything read
- Adding a logit penalty on "wait", "maybe" and "perhaps" is reported to improve Qwen accuracy (r/LocalLLaMA, 2026-09-27). A community finding with no paper and no controlled comparison; if it replicates it is a free intervention on a shipped open-weight model, which is why it is held rather than dropped
- Source-aware verification for MCP agents (HF Blog, 2026-09-29, Multiverse Computing). Getting the source right rather than only the fact — directly on MCP — Model Context Protocol's weakest axis, but a community post with no evaluation returned
New in Wiki
For review.
- Safety Cases (new — concept) — and it immediately makes "containment" the third distinct thing this wiki names with that word, beside Agent Runtime Containment and Eval Environment Containment. Two mechanisms and one argument about a mechanism; the boundaries are undefined and flagged on all three pages
- World Labs (new — entity) —
Models & Productsis a stated absence rather than a placeholder: nothing read names one. Fei-Fei Li is recorded as plain text, not a person page, per the 09-29 precedent - GPT-6.1 Sol (new — model) —
Availabilityis marked as the page's weakest row rather than left to look complete, andCatalogue idisunknownpending the spec-check Action - NVIDIA Kumo Tabular (new — model) — the first model page here outside language, vision, audio and video.
Context windowisunknownrather than dropped: the model reads a table in context, so the row applies and nothing stated it - 4 paper pages: Post-Training Leaves Behavioral Shadows on Unrelated Decisions, Relic: From Multi-Agent Collaboration to Persistent Organizational Capability, Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL, Imprint Reader: From Weight-Update Readout to Behavioral Intervention
Updates
- Agentic Reinforcement Learning: the verifier is a harness too. Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL names why binary test-pass rewards make GRPO blind to code quality — identical advantages for every passing trajectory — and fixes it with sum-preserving advantage redistribution behind an SFT-trained grader, on pre-RL checkpoints of MiMo-V2.6-Flash and Pro. It publishes not one number for any of its three claimed improvements. Captured with it: Alibaba / Qwen AI Lab's QwenGyre, NL2RepoBench 52.5% → 58.5% in 48 steps on Qwen 3.8 2.4T at 700K tokens/rollout, 1.85×/1.78× over Colocate and Async — a one-off mention, and its 2.4T corroborates this wiki's July figure from the model's own team
- MiMo-V2.6-Pro: 310B or 309B? GAGAR states MiMo-V2.6-Flash at 310B total; this wiki holds 309B from third-party reporting. 1B apart, most likely rounding, and not silently reconciled — the paper is effectively the developer's figure against a journalist's. Disclosed under
## Conflicting Reports. The 1.02T figure for Pro is now independently confirmed - Kimi K3: Amazon Bedrock from 2026-09-18, captured +12 days on the Moonshot rotation slot — a surface change, not a release. Recorded with it: no Kimi K4 exists; one July report of Moonshot seeking Blackwell-class capacity, and no name, count, benchmark, price or date from Moonshot
- Astra: it is now the model dots run on, unattended and indefinitely — the first product here where Astra runs outside a request boundary. Also the figure GPT-6.1 Sol's price claim is measured against
- Google DeepMind: feed recovered. Yesterday's
ParseErrorondeepmind.google/blog/rss.xmlis gone — 12 of 12 feedsOK, 100 items,last_ok: 2026-09-29. One bad fetch that resolved on the next run, which is what the 2026-07-26 precedent said to expect