AI Trend Notifier
EN한
← archive

$ cat briefs/daily/2026-09-25.md

2026-09-25

September 25, 2026 (Fri)

5 stories · 2 paper picks · 4 watch items · 4 new pages

**The largest finding today is in this pipeline rather than in the news, and it is recorded first because it nearly put a duplicate page on the wiki.** Yesterday's run captured MentalHealthBench onto two wiki pages and committed its snapshot, and **never wrote a line about it in `log.md`** — which is the file `agents/ingest.md` fixes as the dedup key. So this run could not see it, re-derived the benchmark from scratch, and overwrote the committed snapshot before the check caught it. The snapshot was restored, the duplicate page deleted, and **nothing under `sources/` is modified in today's commit**. **`cloud.google.com` answered first-party and produced the only vendor-read announcement of the run** — the first time a Google announcement has been read as written on this wiki. `www.anthropic.com` answered for the third consecutive run. **Thirteen of fifteen other hosts answer `EGRESS_BLOCKED`**, including `deepmind.google`, `blog.google`, `openai.com`, `x.ai`, `www.meta.com`, `press.un.org` and `blog.arxiv.org`, so most of what follows is assembled from search passes and says so.

+4new pages
[01]

Top Stories

1. Amodei asked the UN Security Council for the antitrust waiver five days after being sued over it (1.59 — the highest-scoring story; the highest score on the page is a Paper Pick at 1.65)

  • The Council's 10228th meeting, Artificial Intelligence and International Security, 2026-09-23, convened by France as September's president and chaired by Jean-Noël Barrot (source)
  • Briefers: Yoshua Bengio (co-chair, UN Independent International Scientific Panel on AI), Sam Altman, Dario Amodei (remotely, per one pass), and Clément Delangue of Hugging Face
  • Amodei restated the three steps of his 2026-09-12 essay: embedded evaluators with "employee-like access" to training pipelines — "desks, badges, company laptops, and a right to publish findings the lab cannot bury" — democratic coordination requiring a narrow US antitrust waiver, and global coordination with China from AI for biological weapons toward "speed limits" on recursive self-improvement, modelled on Cold War SALT treaties
  • Altman publicly endorsed the embedded-evaluator step, said "no level of catastrophic risk is acceptable", and that companies should not train models unless they can make a strong case that those models will stay under human control — a precondition on training, not on release
  • Bengio's four are the only proposals a state imposes rather than labs agreeing with each other: licensing frontier models, liability insurance, incident reporting, shared safety requirements
  • Why it matters: Buist v. Anthropic PBC was filed 2026-09-18 pleading exactly that coordination as a per-se Sherman Act violation. The waiver was asked for again, from the Security Council, five days later — and the endorsement came one day after OpenAI's own assessment principles published criteria without naming Anthropic's engagement, a document that still does not reference what its CEO endorsed
  • Not established: no outcome, resolution or presidential statement; no member-state position of any kind, which for a Security Council session is the substance rather than a detail; Delangue is named in two passes with nothing attributed to him, leaving the one open-weights voice silent on licensing; and whether the waiver was put as a request or restated as background — no pass quotes Amodei on it directly
  • → Frontier Pacing · Embedded Evaluation · AI Governance · Anthropic · OpenAI

2. Live Avatar went generally available, and the model it runs on finally has a price (1.46)

  • Live Avatar is a capability on Gemini 3.8 Live, not a new model, generally available 2026-09-24 in Gemini Enterprise on US and EU endpoints — read first-party on the Google Cloud blog (source)
  • What it adds: video avatars with synchronized lip-syncing, 97 languages with automatic detection, live camera feeds and screen shares processed alongside audio, tool calls running in the background while the conversation continues, and SynthID watermarks on all generated audio and video
  • Custom avatars are allowlist-only, behind a Google Cloud sales contact and enterprise verification; everyone else picks from a curated pre-built library
  • Gemini 3.8 Live Extended Thinking remains in private preview and did not go GA alongside it — so nine days after the sibling carried every published figure, the general-availability model is the one with the capability
  • Pricing has stopped reading unknown: $0.005 per minute of audio input and $0.018 per minute of audio output, roughly $1.38 per hour of conversation. The cell is marked (third-party) — several passes agree, no Google rate was obtained, and a vendor figure outranks it the day one appears
  • Why it matters: no page was created, per the one-page-per-model rule. Treating a feature as a model is how a wiki acquires pages with nothing to compare, and this is the first time that rule has been applied to a capability GA rather than to a series
  • Not established: no price for Live Avatar on top of the base rate; no API model id; no benchmark, latency or frame-rate figure; no avatar count; no stated allowlist criterion; and no consent, likeness or impersonation safeguard beyond SynthID, against a feature whose output is a speaking human likeness
  • → Gemini 3.8 Live · Google DeepMind

3. Daybreak went to a state at war, and the line that makes it defensive is a sentence (1.20 — org weight only; interests.md has no cyber or security row, and none was invented)

  • OpenAI will give the Government of Ukraine access to Daybreak, working with the Ministry of Digital Transformation, to identify software vulnerabilities and develop and test fixes for civilian infrastructure (source)
  • Announced on the sidelines of the UN General Assembly by Dmytro Kushneruk, Consul General of Ukraine in San Francisco, and Sasha Baker, OpenAI's Head of National Security Policy
  • Context given: CERT-UA handled nearly 6,000 cyber incidents in 2025, including attacks on hospital systems, the energy sector and telecommunications
  • Why it matters: every prior Daybreak entry on this wiki is about who may use an offensive-capable model. This is the first about giving that access to a combatant state, and the constraint that makes it defensive — "civilian infrastructure rather than offensive military cyber operations" — is a sentence with no control described anywhere read, against networks in active wartime use
  • Not established: no cost, term or duration — two passes headline it "free" against OpenAI's own "extends access" framing; no model is named; no access control, audit or safeguard beyond the word "authorized"; no seat count; and whether Ukraine is the first government to receive Daybreak access
  • → AI-Enabled Cyberattacks · OpenAI

4. Google is giving Private AI Compute a memory, and keeping the keys on the phone (1.20 — org weight only, same reading as story 3)

  • A persistent server-side memory layer in which data sits in dedicated encrypted storage in the cloud while the cryptographic keys are held exclusively on the user's personal devices; when a model needs stored information, an authenticated end-to-end encrypted channel connects the device to an isolated cloud environment (source)
  • External auditors validated the design for both the initial release and this update, and Google states it has published summaries of the 2025 and 2026 audit reports
  • Why it matters: this is the first announcement this wiki holds that proposes cross-device assistant continuity and on-device privacy are not a trade-off, and it describes the mechanism that every memory feature recorded here has declined to describe
  • Not established: no model is named, and whether any Gemini model uses it yet; no availability date, surface or region — one pass says Google "will bring" it, which reads as forthcoming rather than shipped; the auditors are unnamed and neither summary was read; no retention period, user-visible control or deletion guarantee; no threat model for a lost device; and no stated relationship to Safety Monitoring and Data Retention
  • → Google DeepMind · Safety Monitoring and Data Retention

5. A second xAI release missed inside one week, this one by seven days (1.00)

  • Grok Voice Transcribe 2.0 (new) — Grok Voice Transcribe 2.0, released 2026-09-18, built on the audio foundation model behind Grok Voice (source)
  • $0.10 per hour batch · $0.20 per hour streaming — unchanged from 1.0, against a claim of being "twice as accurate" as 1.0, which another pass renders as "half the errors"
  • Claimed first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard; features are word timestamps, diarization, multichannel support, and automatic language detection with mid-recording switching handled in a single pass
  • Why it matters: the cause is identical to the Grok 4.7 capture at +3 days three days ago, and it is structural — xAI publishes no feed, so nothing it ships reaches state/prefetch.json; the Chinese-lab rotation does not cover it; and x.ai answers EGRESS_BLOCKED, so the non-feed sweep is a search query against a lab whose releases do not surface reliably under its own name. Two misses in one week is a pattern with two instances, and it goes to the W39 lint
  • Not established: no API model id; no word error rate, baseline or dataset behind either accuracy claim; no licence; no audio-length limit; no latency figure; no model card. The leaderboard claim could not be checked against this repo's own snapshots — aa-fetch.py captures the intelligence board, not a speech board
  • → Grok Voice Transcribe 2.0 · xAI
[02]

Paper Picks

Agensh: Scaling Organizational Intelligence to 1,024 Agents — arXiv 2609.26781 (1.65 — the highest score on the page)

  • TL;DR: a multi-agent harness with no central orchestrator — workers gather context, claim and self-assign sub-tasks, act, verify and merge asynchronously over a shared workspace, a message interface and shared context. On the five hardest ProgramBench tasks with GPT-5.6-sol (high), 1 → 128 agents raises the mean final test-pass rate 19.31% → 28.78%; on pandoc alone, 1 → 1,024 agents raises it 33.89% → 55.06%
  • Why read it: it proposes the agent count as a scaling dimension — and it reports no token, dollar or wall-clock figure at all, against a paper whose stated benefit is latency. 1,024 agents on one task is 1,024 times the inference, and the reported quantity is a pass rate
  • It contradicts Harness-Zero: Harness Distillation via Agent-as-Harness from two days earlier without citing it. That paper raised performance by removing the scaffold — 23.3% → 44.3% harness-free against 41.7% with one attached. The field published "the scaffold is the constraint" and "the scaffold should be a thousand times bigger" 48 hours apart, on different benchmarks with different models
  • → Agensh: Scaling Organizational Intelligence to 1,024 Agents

Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? — arXiv 2609.27891 (1.39)

  • TL;DR: SchrodingerRepo instantiates the test repository at evaluation time — problem-statement reconstruction, namespace remapping, intra-file layout reordering, functionality-preserving rewriting — preserving executable behaviour and destroying familiarity. On SWE-bench Verified and SWE-QA, removing repository cues consistently degrades performance and substantially increases interaction cost, the extra cost landing on exploration and localization
  • Why read it: it is the first testable version of the contamination objection this wiki holds, and it lands on a live gap — Claude Opus 5.5 shipped on 2026-09-22 with no SWE-bench-family row at all, and GPT-6 Sol and GPT-6 Luna shipped the same day with DeepSWE v1.1 as their only benchmark. This supplies a reason to move off SWE-bench that is not "we score worse on it" — while also undermining the replacement, since DeepSWE is built from the same kind of repository
  • Its own record is entirely directional: no percentage-point drop, no cost multiplier, no model named, no ablation across the four transformation levels, and no harness named — the standing gap, reproduced by a paper about measurement hygiene
  • → Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
[03]

Watch

  • A capture reached the wiki and the snapshot but never reached log.md, and log.md is the dedup key. Yesterday's entry says "New sources: 6" and names five. This run re-derived the missing one from scratch and briefly overwrote a committed, append-only snapshot before the check caught it. Dedup should key on sources/ and the wiki pages as well as on the log — carried to the W39 lint → Eval Harness Configuration
  • Meta Connect 2026 was an hour about Muse with no model in it (0.77 — interests.md weights product launches at 0.7, and the rule was applied rather than argued around): Meta VR Glasses $1,299 in spring, the Muse Charm pendant carrying the agent standalone in December with no announced price, Ray-Ban Meta Audio at $349 on 2026-10-13, plus Walmart and Instacart integrations and Muse Space. Zuckerberg: Muse is "free for a huge number of tokens" with profit expected from "a small fee from transactions". No model version, parameter count, context window, benchmark or API for Muse appears anywhere → Meta AI
  • A closed model's hidden reasoning was read out through an ordinary API feature (1.39 — tied with the second Paper Pick; it is here rather than there because the other two pair on one finding): Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models registers a simple custom tool on a standard API feature to induce externalized reasoning, validates against native CoT on open models first, then extends to closed ones including GPT-6 Astra — extracted reasoning matches native performance. It is an unintended disclosure channel in the week Claude Opus 5.5 shipped "preserved thinking" as an anti-distillation safeguard; nothing read connects the two, no vendor has responded, and the API feature is not named, so the result is neither reproducible nor mitigable as published
  • alignment.anthropic.com unreachable for the eighth consecutive run, so the article list still cannot be checked against sources/ — the gap that cost this wiki two posts at +73 and +38 days. And the X-account sweep returned no dated high-signal item at all this run, the first time it has come back wholly empty → Anthropic
[04]

New in Wiki

No new entity or concept page was created today, and two were deliberately not written: people/bengio, wanted by two documents, which the wait rule says is too thin to build; and concepts/mentalhealthbench, which this run wrote and then deleted on discovering the subject was already held on OpenAI and Eval Harness Configuration.

[05]

Updates

  • Frontier Pacing: new dated section — the antitrust waiver asked for again at the Security Council, five days after the Sherman Act complaint
  • Embedded Evaluation: Altman's endorsement attaches the two halves this page has tracked separately, and Amodei's "right to publish findings the lab cannot bury" states the publication right both founding documents leave unrecorded
  • AI Governance: the session's proposals split by who imposes them — three lab mechanisms requiring labs to agree, Bengio's four requiring a state
  • Eval Harness Configuration: the repository is a variable too, and the two harness papers that contradict each other
  • Agents (LLM Agents) · Mechanistic Interpretability · AI-Enabled Cyberattacks: the three new papers and the Daybreak extension
  • Gemini 3.8 Live: Pricing filled from reporting after nine days at unknown; Live Avatar GA recorded on the model rather than on a new page
  • Grok 4.7: max output, open on 2026-09-24, now has a claim and not a figure — one pass says "no text output limit", recorded and not adopted as a spec value
  • Anthropic · OpenAI · Hugging Face · Google DeepMind · Meta AI · xAI · index.md