AI Trend Notifier
EN
← archive

$ cat briefs/daily/2026-09-24.md

2026-09-24

September 24, 2026 (Thu)

5 stories · 2 paper picks · 4 watch items · 6 new pages

**Four of today's six captures are consequences of pages this wiki already held, and one is a model this pipeline twice reported as unreleased.** Grok 4.7 shipped on 2026-09-21; the runs of 09-22 and 09-23 both wrote that it had not. OpenAI's assessment principles answer the mechanism [[concepts/embedded-evaluation]] was created for four days earlier. China's regulator opened an investigation on the strength of Anthropic's 09-10 threat report. And two papers measured [[models/jev]] nine days after [[entities/typesafe]]'s page recorded that no independent measurement existed. **`www.anthropic.com` answered first-party for the second consecutive run**, so the enzyme story below is read as written. Eight of nine other hosts answer `EGRESS_BLOCKED` — `openai.com`, `deepmind.google`, `x.ai`, `triviumchina.com`, `techcrunch.com`, `alignment.anthropic.com` and two trade outlets — so every other story here is a search extract with its identity fixed by the publisher's own feed.

+6new pages
[01]

Top Stories

1. Grok 4.7 shipped three days ago and this pipeline said twice that it had not (1.20 — led here under interests.md's "new frontier model announcement" tracked signal, not on score)

  • Grok 4.7, released 2026-09-21: 2.1T parameters against Grok 4.6's 1.5T, 500K context, four effort levels with high the default (source)
  • The price did not move: $2/M input · $0.50/M cached · $6/M output at prompts up to 200K, identical to Grok 4.6, doubling above that threshold. Available in the Grok API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare
  • Vendor benchmarks: CursorBench 4.0 40.4% → 46.3%, EEBench 53.0% → 64.0%, Harvey Legal Agent Benchmark 15.8% → 19.6% — the last reported against GPT-5.6 Sol 2.5% and Fable 5.1 6.7%
  • It arrived nine days past the third expired window and thirty past the first, ending the delay this wiki has been counting since 2026-08-22
  • Why it matters: the comparison set is GPT-5.6 Sol and Fable 5.1, both superseded the next day, so no figure connects this release to the 09-22 wave — and a 7.8× ratio on a benchmark where the whole field scores under 20% is a spread, not an ordering
  • The miss is ours and it has a cause: xAI publishes no feed, so it never enters state/prefetch.json; the lab-named rotation query returned nothing; the release surfaced only through a version-named query, the third instance of that pattern since 09-20
  • Not established: no max output, no licence, no model card, no safety evaluation, no independent measurement — and whether the RL length-penalty defect Musk named on 09-11 was fixed, since the stated "longer RL pass" is adjacent to that diagnosis and is not offered as a fix for it
  • Grok 4.7 · xAI · Eval Harness Configuration

2. OpenAI published the rules for how it gets evaluated — and named no evaluator (1.59 — the day's highest-scoring story)

  • Priorities and principles for effective third party assessments, 2026-09-22, framed by OpenAI as "part of our efforts to pace the frontier" (source)
  • Four priority areas: its safety cases examined from training through internal use to release; whether critical safeguards hold in realistic conditions; whether its Preparedness Framework capability evaluations for chemical and biological risk, cyber attack and AI self-improvement "actually measure what they claim"; and independent investigation of misalignment incidents
  • Seven principles, among them pre-registered claims against a mutually agreed scope, proportionate access falling back to a designated company representative where direct access is impractical, and actionable findings with time for remediation
  • Lama Ahmad, who leads OpenAI's work with outside safety experts, says the company previously brought third parties in shortly before launch
  • Why it matters: Anthropic's Accenture deal four days earlier named a counterparty, a price and a scope with no published criteria. OpenAI published criteria with no counterparty, no price, no funding arrangement and no start date. The two halves of a working arrangement now exist at two different companies, and nothing read connects either to the other, or either to the AI Evaluator Forum letter published the same day as Anthropic's
  • Not established: the AEF's three independence criteria ask who the assessor may be; this document addresses how an assessment is run, so it answers none of them. The publication right remains unrecorded by both labs — "balancing transparency with confidentiality" does not say who decides. And still no Preparedness Framework classification for GPT-6 Sol or Luna, now two days old
  • Embedded Evaluation · OpenAI · AI Evaluator Forum (AEF)

3. 950 Claude agents read DNA for 21 hours and came back with an enzyme system nobody had characterised (1.30 — org weight only; interests.md has no row for AI-for-science, and none was invented)

  • ART — "array-associated reverse transcriptases" — found mainly in bacteriophages, comprising a reverse transcriptase, a partner gene beside it, and "a long array of evenly spaced DNA repeat sequences", the CRISPR-like structure the name comes from. Read first-party (source)
  • How: approximately 950 Claude agents, 21 hours, 210 million tokens against a DNA sequence database. The only human direction was the opening prompt. Feng Zhang (MIT, Broad Institute) is quoted endorsing it
  • Alongside it: Claude Science optimising 30+ open-source biomolecular models in under four weeks for roughly 4× speedup, 14 of 15 targets given binders validated in Adaptyv Bio and Twist Bioscience wet labs, and a competition putting 5,000+ designs through Adaptyv's automated lab with up to $1M in Claude credits
  • Why it matters: every autonomous-research claim this wiki holds is a score — R&D Automation Index at 26% on AL4, the harness papers, the RSI argument on Frontier Pacing. This is the first whose output is a physical object in a database, and therefore checkable by a route none of the others are
  • Not established, and it is the important half: the function of ART is unknown and Anthropic says so. A structural resemblance to CRISPR is not a functional claim. No peer review, no model version for the agents, and the agent, hour and token counts are Anthropic's own. The only experimental result is that the array is expressed as a set of distinct short RNAs
  • Anthropic · R&D Automation Index

4. A US lab's threat telemetry became a Chinese regulator's investigation — and the charge runs the other way (1.20)

  • China's Cyberspace Administration is reported to be probing DeepSeek and Moonshot AI over Anthropic's claim that both covertly routed user requests through Claude (source)
  • Volumes this wiki could not obtain on 09-10 now appear in reporting: Moonshot over 23 million exchanges across May–July, DeepSeek 12.1 million in a 14-day July window, from a 154-page report naming seven labs
  • Regulators are examining whether sensitive Chinese police, military and state-linked data reached a US system. The cited example is a user Anthropic assessed as likely PLA-affiliated asking Kimi to trace a person across hundreds of police cameras in Chengdu, with the request and footage allegedly passed to Claude without telling the user
  • Officials have visited both firms to question executives and staff. The probe is ongoing, with no penalties determined and no sanctions announced as of 09-22
  • Why it matters: Anthropic alleged value taken from a US model. The investigation is about data going to one. The same conduct grounds two opposite complaints in two jurisdictions, on evidence supplied by a commercial competitor that no regulator has verified in anything read
  • Not established: no CAC statement, notice or legal instrument — the probe is reported, not published. Which regulation is unnamed, neither lab has responded, and the other five named labs are not reported as under investigation
  • AI Governance · DeepSeek · Moonshot AI

5. Google shipped two speech models, 130 languages, and not one number to check (1.20)

  • Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, announced 2026-09-23, both rolling out same-day in the Gemini API and Google AI Studio (source)
  • Documented for both: 8,192 text input tokens, 130 supported languages, voice design and replication, and a SynthID watermark on every generated clip
  • Flash TTS is positioned for creative direction and character design; Flash-Lite for dubbing, bulk audio and voice agents — the first Google speech model on this wiki aimed at the agent stack rather than media production
  • Why it matters: two pages exist because this wiki writes one page per model, and the consequence is recorded on both — the published specifications are identical, so the only thing separating them is what Google says they are for
  • Not established: no price for either model, which leaves the cost-efficiency claim that justifies the split with nothing under it. No benchmark, MOS score or listener-preference figure of any kind, so "most expressive yet" is the entire quality claim. No latency, no throughput, no voice count, no licence — and voice replication is named as a capability with no consent or likeness safeguard beyond SynthID
  • Gemini 3.8 Flash TTS · Gemini 3.8 Flash-Lite TTS · Google DeepMind
[02]

Paper Picks

"JEV-as-a-Judge: Accept When Confident, Escalate When Unsure" — arXiv 2609.26550 (1.65 — the day's top score)

  • TL;DR: a decision-only judge against sixteen generative and reward-model judges under blinded human adjudication, landing within three percentage points of the strongest comparator on ordinary preference and evidence-grounded factuality at 0.36% of that comparator's fee. Its gap is concentrated in low-confidence decisions, so a frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy
  • Why read it: it is the first evaluation of Jev this wiki holds that is not the vendor's — and it splits the vendor's two claims. The accuracy figure survives an outside harness at the same three points. The cost multiplier does not: 0.36% of a fee is ~278×, against TypeSafe's ~4,000×, a 14× spread between figures that are not measuring the same unit
  • The caveat is the word the result rests on: neither snapshot entry carries an author list or affiliation, so independence is unestablished and is not asserted
  • JEV-as-a-Judge: Accept When Confident, Escalate When Unsure · companion Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents, which moves five routing decisions off the autoregressive model inside an agent's memory: LoCoMo 0.777, +11.0% relative, construction 158 s at 6.6×, latency 0.93 s at −36.7%

"The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks" — arXiv 2609.25804 (1.65)

  • TL;DR: Taste-Bench stops measuring whether an agent finished and starts measuring the decision forks inside the run — mined automatically from parallel attempts at the same task and from detours inside one trajectory, with no human annotation. The best frontier model answers 59.7% correctly
  • Why read it: a larger reasoning budget does not improve accuracy, which is a negative result against the direction Test-Time Compute (Inference-Time Compute Scaling) has been tracking all year. And taste distils: a teacher that saw the outcome trains a student that did not, improving end-to-end success on held-out SWE-bench Pro
  • Not established: no model is named for the 59.7%, no harness is named at all, and the distillation result carries no figure — it is directional only. Forks whose deciding evidence appears later in the trajectory are reported as much harder, and nothing read separates poor taste from an unanswerable question
  • The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
[03]

Watch

  • alignment.anthropic.com unreachable for a seventh consecutive run. The article list still cannot be checked against sources/. Search returns only items already held — this is a source cited 24 times by this wiki that nothing has fetched in a week
  • Reasoning (derived) has read no on every row of every Artificial Analysis snapshot since 2026-09-03, against 153 rows marked yes on 08-02. Found on 09-22, still not fixed, and deliberately so: the host is blocked from this sandbox, so any selector change would be a guess pushed into a scraper that writes committed snapshots. Carried to the W39 lint as an action item
  • MiniMax M3 Pro held out for a thirteenth consecutive day — no announcement, API string, model card or technical report. One pass still dates it "targeting Q3 2026", and Q3 ends in six days
  • A wikilink inside a markdown table cell failed silently today. […](/wiki/concepts/preparedness-framework) needs its pipe escaped to render in a table, and the escape makes the link target concepts/preparedness-framework\removing the edge from the graph while still rendering as a link. Caught by the pre-brief sweep and repaired. Worth watching for because the page looks correct
[04]

New in Wiki

[05]

Updates

  • TypeSafe AI and Jev: the ## Strategic Position sentence recording that no independent measurement existed is now nine days old and partly answered. A ## Conflicting Reports entry was added for the 4,000× against 278× cost discrepancy
  • Embedded Evaluation: ## State of the Art advanced to 2026-09-24 and now carries two documents from two labs sharing not one element, plus a new Open Problem — three publications in five days on one question, none citing another
  • Eval Harness Configuration: a new entry on the first cross-vendor score spanning the 09-22 wave, published by one of the vendors. The reproducibility of the released data is the load-bearing part, not the leaderboard
  • Model Routing: routing has moved inside the system — a cascade routed by the cheap model's own uncertainty, and five control decisions taken off the critical path of an agent's memory
  • AI Governance: ## State of the Art advanced to 2026-09-24 with a mechanism not previously recorded — a private firm's threat telemetry as the evidentiary basis for a state investigation of foreign companies
  • xAI, OpenAI, Anthropic, DeepSeek, Moonshot AI, Google DeepMind, index.md