$ cat briefs/daily/2026-10-04.md
2026-10-04
October 4, 2026 (Sun)
5 stories · 3 paper picks · 5 new pages · a two-day gap closed · an executive order that renames the subject
**No run produced anything on 2026-10-03**, so this is a fallback run covering two days. The practical consequence is in 📄: the HuggingFace Daily Papers snapshots for **10-02** (26 ids the 10-02 run recorded as owed) and **10-03** (which no run read) were both consumed deliberately rather than skipped past, and two of today's three paper picks come from them. **The day's most consequential item ranks fourth, and the brief publishes the reason rather than reordering.** The US executive branch has, by executive order, stopped using the term "Artificial Intelligence". It scores **1.10** because `interests.md` still has **no governance row** — a gap flagged on every run since 2026-09-29 — so a Presidential document is weighted as a neutral 1.0 topic from an unweighted org. See 📊. Two reachability facts, both changes rather than repeats: **`alignment.anthropic.com` answered for the first time in sixteen runs**, and **`mistral.ai` answered for the first time**. Neither produced new material, and that is recorded as a result.
Top Stories
1. A European lab shipped Apache-2.0 weights where only 3.46B of 78.1B parameters run at a time — and the efficiency claim holds for maths, not code
- Kolibri-1 from Aleph Alpha, released 2026-10-03: 78.1B total / 3.46B active, 384 experts with 6 active, 50 layers, Apache 2.0, full weights on Hugging Face (source)
- The stated claim is that it "matches models with up to four times its active parameter count". Its own table: AIME 2025 96.9 and GPQA Diamond 84.3 against Nemotron 3 Super 120B's 91.7 and 78.0 — but HumanEval+ 92.7 against 94.7, and BFCL v4 61.4 against 61.0, a tie
- So the claim is carried by maths and science and not by code or tool use, which matches the training mix: ~14% code against ~62% English and 21.3% German
- The sovereignty argument is unusually checkable: 4.3 trillion German tokens of a 20-trillion-token three-stage run — the highest disclosed non-English share of any model this wiki holds
- Native context is disputed and not resolved: the first-party post says 16,384 native extended to 1,048,576; secondary coverage says 262,144 native. The first-party figure is carried and the disagreement is recorded on the page, because a 16× gap changes what the extrapolation is doing
- Why it matters:
Pricingis unknown because there is no list price and no hosted endpoint — the distribution route is download the weights. An open model whose pitch is active-parameter cost rather than leaderboard position is competing on who can afford to run it, which is the axis Open-Weights Policy Fight tracks and the one frontier releases do not touch - → Kolibri-1 · Aleph Alpha · Open-Weights Policy Fight
2. An agent harness that spent a year declining to implement MCP shipped it in its 1.0
- Earendil's Pi v1.0.0, released 2026-10-01, carries MCP natively through Codemode, a harness-side JavaScript sandbox the model uses to orchestrate tool calls (source)
- The security specifics are the substance: RFC 9207
isschecks, credentials per server, and step-up sign-in that "keeps granted scopes" — the two named mitigations for the confused-deputy and token-reuse problems sitting in MCP — Model Context Protocol's## Open Problems, reported as shipped by an implementation - Codemode uses ~40% fewer prompt tokens than its previous version, with "errors that tell the model how to recover"
- Two widely repeated claims are not adopted: "Pi Durable" appears nowhere in
the v1.0.0 release notes despite leading the Latent Space AINews headline, and the
#1 on Hacker News ranking could not be checked because
news.ycombinator.comwas not fetched - Why it matters: MCP — Model Context Protocol has tracked the spec and its weaknesses but never a holdout adopting it. A client whose prompt-token cost went down on adding protocol support is evidence against the assumption that MCP generality is paid for in context — though the figure is a release note's self-report against its own prior version, not a comparison with any other client
- → MCP — Model Context Protocol
3. A 13-million-line machine-checked proof of Fermat's Last Theorem, built in 11 days — published 30 days ago and read today
- Formalizing Fermat's Last Theorem, Anthropic Research, 2026-09-04, primary author Tianyi Peng with Kevin Buzzard consulted (source)
- 30,300 theorems proved, 29,500 in the final proof, 13 million lines of Lean, 11 days, about six billion output tokens from "a general-purpose internal research model roughly comparable to Claude Fable 5.1"
- The platform is not Anthropic's: Prove2Me, "an open collaborative platform for formalizing mathematics designed by Tianyi Peng and his collaborators at Columbia University", holding a DAG of theorem statements so agents could work in parallel
- Disclosed limits: the proof is "likely much longer than it needs to be", and early agents "quickly lost track of the project's state and stopped collaborating effectively" — contributing ~7% of non-boilerplate lines through failed iterations
- Why it matters: FLT was proved by Wiles in 1994, so there is no novelty claim here to contest — which makes it the cleanest case on AI for Mathematics of the generation/verification split that page is built around, falling entirely on the verification side. A 13-million-line artefact anyone can re-check beats a result only the authors can vouch for, and the 13-million figure is itself the "grinds through calculations rather than building tools" tendency made countable
- Captured +30 days because it sits on
anthropic.com/research, the path the 10-02 run discovered and recorded this post as a title only - → AI for Mathematics · Anthropic
4. The US executive branch will no longer say "Artificial Intelligence"
- Executive Order 14434, Inaugurating the Era of Super Intelligence, signed 2026-09-29, published 2026-10-02 at 91 FR 63129, directs agencies to use "Super Intelligence" and "SI" in place of "Artificial Intelligence" and "AI" (source)
- The operative content is terminology and nothing else. § 2 reaches correspondence, public communications, websites, reports and policy documents; § 2(b) exempts existing regulations, contracts and grants. No agency gains or loses authority, no evaluation or reporting duty is created, and no model, lab or capability threshold is named anywhere in the order
- § 3(a) borrows the existing legal scope — "SI" is defined as what 15 U.S.C. § 9401(3) already calls artificial intelligence — so on the day of signing only the label moves
- § 3(b) is the part with teeth: the Assistant to the President for Science and Technology has 60 days (to roughly 2026-11-28) to submit proposed legislative language for a federal definition, explicitly assessing whether it should modify, expand upon, or supersede the statutory definition
- One clause goes past renaming: the executive branch "will not acknowledge the usage of 'Artificial Intelligence' and 'AI' in any applicable setting" — and "applicable setting" is undefined in the order
- Why it matters: most US federal AI obligations are keyed to § 9401(3), so an instruction to go rewrite that definition is an instruction that reaches all of them. The rename is the visible part; the 60-day proposal is the one to watch, and the order requires no publication of it
- → AI Governance
5. Anthropic will spend $100M training 10,000 engineers inside its own customers
- Claude Frontier Academy, 2026-10-02: a Frontier Deployed Engineer Residency targeting 10,000 engineers by end of 2027, built as multi-day in-person training, a simulated deployment exercise and a 12-week residency running real Claude projects at the engineer's own organisation (source)
- First cohorts: Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley, Novo Nordisk
- Stated rationale: "A small group of deeply skilled people drives an outsized share of what AI delivers." Steve Corfield: "No AI company has invested in developing that talent inside its customers and partners at this depth."
- The one outcome figure is unbaselined. Commonwealth Bank's "Our teams have produced up to 3x more code changes in the past year" names no prior rate and does not attribute the change to Claude, so it is recorded as a quotation and not as a result
- Why it matters: the constraint being bought here is deployment capacity at the customer, not model capability — which is a claim that the bottleneck on enterprise adoption is people who know how to wire the model in. Also not stated: per-engineer cost, cohort size, start date or selection criteria
- → Anthropic
Paper Picks
Three picks, two of them from snapshots no run had read. arxiv.org is blocked from
this pipeline, so all three rest on the HuggingFace Daily Papers abstracts and
none has a stated author or affiliation — none is guessed.
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents — arXiv 2609.39102
- TL;DR: when a proposer generates questions and a solver answers them under joint optimisation, the two "increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness" — and it worsens over successive rounds while the training curve keeps rising
- False-agreement mass falls 6.1% → 3.0% (Qwen3.5-4B) and 8.8% → 3.7% (9B) under CrossFit, which partitions source documents and scores each half with a solver trained only on the other. +8.8 / +8.4 points across seven search benchmarks over coupled self-evolution
- Why read it: the usual remedy for reward hacking — a better reward model — is unavailable when the system generates the objective it is scored against. The fix is structural: remove the information path, don't verify harder. MSV, the verify-harder option, costs six extra generations per candidate and still leaves "substantial residual co-cheating". The number to keep is that baseline co-cheating is worse at 9B (8.8%) than at 4B (6.1%)
- → False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents — arXiv 2609.39982
- TL;DR: sample candidate shell actions and verify them before execution, leaving generator and harness untouched. TerminalBench-Lite Pass@1 50.00% → 68.03% with 8 sampled actions and a GPT-5.6 Sol verifier
- Why read it: the conditional is the finding, not the 18 points. More sampling "yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator" — the better actions were already in the distribution and what was missing was discrimination. No cost accounting is given for calling a frontier model on every action
- → Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics — arXiv 2609.35259
- TL;DR: vary rollout policy, token-level KL direction and learning rate independently and rollout policy "does not necessarily play a central role" — KL direction shapes performance and coverage, learning rate governs forgetting and update sparsity
- Why read it: it reassigns all three benefits usually credited to on-policy training, and explains why the literature disagrees — forward KL is robust to rollout policy, reverse KL is not, so confounded comparisons could honestly report either answer. No numeric figure of any kind appears in anything read, so every finding is directional
- → On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Watch
- The APST legislative proposal due ~2026-11-28, not the rename. EO 14434 § 3(b) asks whether the new definition should supersede 15 U.S.C. § 9401(3); the order states no publication requirement, so this may only become visible if Congress receives it (source)
- Both Sunday leaderboard snapshots are missing.
lmarena-2026-10-04.mdandartificial-analysis-2026-10-04.mddo not exist. LMArena's last capture that parsed is stilllmarena-2026-09-20.md— 14 days old, against 15 wiki pages citing an LMArena snapshot; Artificial Analysis last captured 2026-09-27. No scraper was run from here, per standing policy. Carried to lint 2o - A rank-8 LoRA at one early layer took Qwen3-8B from 15.5% to 99% on 24-line reference chains, with all model weights frozen — Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It (arXiv 2609.36585, HF Daily 2026-10-04, 62 upvotes). Thirteen base models follow only 1.4–3.6 lines by default, so the claim is that "default answers understate the computation accessible through a tiny edit" (source). Not written up as a page this run
- OpenAI published a GPT-6 family model guide on 2026-10-02 naming GPT-6 Astra,
GPT-6 Sol and GPT-6 Luna with selection guidance by task difficulty
(
https://openai.com/index/practical-guide-building-gpt-6).openai.comanswered HTTP 403, so this was not read first-party and no figure from it is adopted — it is developer guidance rather than a release, and no wiki page was changed for it - Two hosts became reachable and neither produced material.
alignment.anthropic.comanswered after fifteen consecutiveconnect_rejectedruns — its full index was read and all 84 listed articles are already held here.mistral.aianswered for the first time; its newest post is 2026-09-28, already held. Both are recorded as network-policy changes, which is what the operating prompt asks for
New in Wiki
- Aleph Alpha (new — entity page, created on the Kolibri release;
Key Peopleis unknown and the lab's founding, funding and earlier model lines were not established by anything read. Needs review) - Kolibri-1 (new — model page; full Spec row set emitted,
PricingandCatalogue idunknown) - False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents (new)
- Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents (new)
- On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics (new)
No entity page was created for Earendil. It is the first appearance here, and the fact that matters is MCP adoption rather than the company, so Pi 1.0 is recorded on MCP — Model Context Protocol under the one-off-mention rule. Flagged because it is a judgement call, not a rule the schema settles.
Updates
- Scores, so the ranking is auditable. Kolibri-1 1.90 (base 1.2 first-party blog × open-weights 1.0 × unweighted org 1.0, +0.4 weights released, +0.3 new entity page). Pi 1.0 1.80 (1.2 × agents/MCP 1.5). Fermat 1.46 (1.2 × neutral 1.0 × Anthropic 1.3, −0.1 already-rich). EO 14434 1.10 (1.2 × neutral 1.0 × unweighted 1.0, −0.1 already-rich). Frontier Academy 0.99 (1.2 × business deals 0.7 × Anthropic 1.3, −0.1)
- A Presidential document ranks fourth, and this is the fifth consecutive run where
a missing
interests.mdrow changed an order. There is no governance row, so EO 14434 took a neutral 1.0 topic weight, and no org weight exists for the US government. The same mechanism put NVIDIA, AMD, DeepSeek and Ai2 at neutral on the four previous runs. The score is printed above so the omission is auditable rather than invisible; the still-missing rows now number three — governance/security, science, and labour economics - One item was scored and then placed below its score on judgement. What do you
want from AI? (Anthropic, 2026-09-29, read first-party today) scores
1.46, equal to Fermat, but it is a recruitment notice for a study running
2026-09-29 to 2026-10-06, not a result.
interests.mdhas no rule that discounts an announced-but-unrun study, so the demotion is a judgement and is recorded as one. Its substance: ~15-minute interviews conducted by Anthropic Interviewer, an AI; a prior December 2025 study of 81,000 people; and Anthropic's own limitation that "Participants in this study will all be people who use Claude, which is not a representative sample of the public" interests.mdtracked signals: one fired. "Consensus-overturning results" — On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics against the on-policy premise. Not fired: no new frontier model announcement (Kolibri-1 is open-weight, not frontier-class, and no frontier comparison is published); "Changes to agents/MCP standards" — Pi 1.0 adopts the standard, it does not change it; and no interest-person signal for the sixth consecutive run- Updated: AI Governance, Open-Weights Policy Fight,
MCP — Model Context Protocol, Agents (LLM Agents), Agentic Reinforcement Learning,
Test-Time Compute (Inference-Time Compute Scaling), AI for Mathematics,
Anthropic,
index.md— 9 - A scaling result was recorded without a page. Scaling Laws for Looped Mixture of Experts (arXiv 2609.40316, HF Daily 2026-10-02, 15 upvotes) is filed on Test-Time Compute (Inference-Time Compute Scaling): ~3× active-parameter efficiency from sparsity, ~2× total-parameter efficiency from recurrence, and a looped MoE matching a ~2× larger non-looped MoE at matched compute. It is not treated as evidence for Kolibri-1's four-times claim — different model, different measurement, and one of them is a vendor's own table
- Three pre-existing figure/unit line wraps were fixed in the English. The
pre-brief sweep found four splits of a figure from its unit; two were introduced
today and two were verified pre-existing at
HEADon Agents (LLM Agents) (**+10.82\npoints**) and Test-Time Compute (Inference-Time Compute Scaling) (**40,115\nhours**). Both split a bold span across lines, which is the**$852\nbillion**failure the operating prompt names. Fixed in the canonical rather than worked around in the Korean, since both files are retranslated today anyway placeholder-check.py --fixtook down 0 markers;--verdictsreturns exists 0 · build 0 · wait 2 · demote 0, unchanged for six runs. Both waits are demand-1:Llama(demand 1 · supply 91),scaling-laws(demand 1 · supply 10). Nothing to act on