$ cat briefs/daily/2026-09-14.md
2026-09-14
September 14, 2026 (Mon)
1 story · 2 paper picks · 3 watch items · 3 new pages
**One capture, four days late, and it is the one that matters.** Anthropic's Frontier Red Team published measurements of what models can do in intelligence targeting and conventional weapons development — not whether they will agree to, but how well they do it — and the headline number is a photo-geolocation median error roughly a quarter of what the top 0.01% of GeoGuessr players achieve. It surfaced only because the Alignment Science sweep checks a blog this sandbox cannot fetch. Otherwise a quiet Monday: the papers carry the rest of the day, the Chinese-lab rotation returned nothing, and Grok 4.7 is four days past its third window with no new date.
Top Stories
1. Anthropic measured what models can do with a drone and a photograph — and published the number against a human expert baseline (2.33)
- 2026-09-10, Measuring AI capabilities in intelligence targeting and conventional weapons, from Anthropic's Frontier Red Team. Captured today at day +4 (source)
- Two domains, framed as understudied next to the standing focus on cyber, biological and nuclear risk: tactical intelligence targeting — "finding where people are based on fragmentary information" — and conventional weapons development — "engineering drones to strike a moving target" (2 passes each)
- Targeting runs three evals (2 passes, same list both times): identity correlation on synthetic multi-platform social data, photo geolocation against human GeoGuessr baselines, and text geolocation with a sandboxed search tool
- The geolocation figure is the one with numbers behind it. Across a stated 6,000 photos, Mythos Preview reports a median distance error of 37.0 km, against Champion Division GeoGuessr players — stated as the top 0.01% of that player base — at 151 km. Roughly a quarter of the elite human error
- Weapons runs three tasks in simulation: terminal guidance to a vehicle, payload drop within a grenade-like radius, and GPS-denied navigation — scored on the share of flights whose median miss lands within five metres, "the approximate lethal radius of a grenade", and on median miss distance
- Kimi K3 is the only non-Anthropic model any pass gives figures for: an 83% hit rate with a 0.4 m median miss — a lower hit rate than Claude Sonnet 5, but a better miss distance when it lands. PRC open-weights models tested were "behind the frontier, but also showed concerning ability to identify and target adversaries, and improve weapon performance" (2 passes, near-verbatim)
- Why it matters: every safeguard this wiki records for Anthropic — refusal training, classifiers, account termination — is attached to Anthropic's platform, and a capability number is the only one of them that still means anything after weights are public. That makes this the missing input to Open-Weights Policy Fight rather than another safety announcement — and it is sharpened by the fact that the post answers a risk that survives weight release with a control that does not: the stated mitigation is new on-platform classifiers, with no name, date, precision figure or coverage statement in any pass
- Read with the caveats attached, because they are large. Nothing was read first-party —
www.anthropic.comand all five secondary outlets attempted answerEGRESS_BLOCKED, so every figure above is a search extract carrying a pass count. A second geolocation figure of 47.2 km is given to Claude Opus 5 by one pass and Mythos 5 by another and is left unresolved on the page. A "first public systematic evaluation framework" priority claim is one pass and not adopted - It is not the same document as the same-day threat report, which coverage conflates with it constantly — that one is held from the 09-13 run and its findings stay on AI-Enabled Cyberattacks
- → Military and Intelligence Capability Evals (new)
Paper Picks
Beyond Solver Verdicts: Generative Reward Models for Autoformalization — arXiv 2609.11085 (1.69)
- TL;DR: a solver saying proved does not mean it proved your statement. The paper names Verdict-Preserving-Unfaithfulness — an incorrect encoding that "executes successfully and matches the expected verdict" — and proves that verdict-only structural heuristics are bounded to chance-level detection on such traces. Generative Verification (GenV) distils an offline Z3-equivalence oracle into a reference-free continuous score by reusing the model's own vocabulary space: 0.961 AUROC, +11.3 points downstream on agentic test-time compute allocation
- Why read it: yesterday's brief carried 25 Fields Medallists arguing that AI mathematics is announced faster than it can be scrutinised. This says the automated half of that scrutiny has a provable blind spot, and that the cheap patch for it is provably no better than a coin flip. Nothing read connects the two — the adjacency is this wiki's
- Z3, not Lean 4, so it does not yet reach the setting AI for Mathematics actually operates in; no baseline AUROC and no dataset named
- → Beyond Solver Verdicts: Generative Reward Models for Autoformalization (new)
Memory as Plans: World-Action Modeling with Memory-Grounded Planning — arXiv 2609.11561 (1.95)
- TL;DR: MaP-WAM spends long multimodal history as planning-time evidence rather than re-feeding it to the executor, keeping executor context length fixed while a World-Action-Progress model runs each plan over an unknown duration, jointly predicting action chunks and progress. 83.3% RMBench, 78.0% on real-robot tasks, with latency "approximately constant" as history grows
- Why read it: it is the same pressure Agents (LLM Agents) records for long-horizon software agents — context growth degrading the loop that runs most often — answered by moving the growth to the loop that runs least often
- It scores above the pick above it and is published second anyway. Base 1.3 × agents 1.5 = 1.95, against GenV's 1.3 × 1.3 = 1.69, because
interests.mdweights agents highest and Embodied Agents claims that weight. The ordering here follows the running story instead, and the discrepancy is stated rather than hidden — the same way 09-04 handled K2 Horizon and 09-06 handled Z.ai. The constant-latency claim is the deployable property and the abstract supports it with an adverb: no baselines, no task counts, no hardware named - → Memory as Plans: World-Action Modeling with Memory-Grounded Planning (new)
Watch
- Simon Willison, So you want to use OpenRouter? (2026-09-11) — unread,
simonwillison.netanswersEGRESS_BLOCKED, so this is the headline and nothing more. Flagged because this repo's ownCLAUDE.mddocuments at length that OpenRouter's headline price is whichever provider is routing at that instant — GLM-5.2 served by 33 providers between $0.72 and $2.31 on input — andspec-check.pywas rewritten to compare against the first-party endpoint because of it. A practitioner write-up of the same routing behaviour is worth reading the first time the host is reachable - xAI: Grok 4.7 is four days past its third expired window with no new date. Two passes confirm no launch page, API model id, price, model card, context window or benchmark table; Musk's 09-11 "needs a few more days to cook" is still the last statement. No page was changed — a fourth day of nothing is not a new fact, and it is recorded here so that "nothing happened" stays distinguishable from "nobody looked"
- The Chinese-lab rotation returned nothing, and both labs it checked appear in today's top story anyway. Z.ai and MiniMax produced no new artefact — GLM-5.5 remains a rumour with no model card, M3 Pro a standing watch item with no announcement or model string, and the leaked MiniMax successor is held out for a third day. But GLM-5.2 is named in Anthropic's weapons evals, which is the first time on this wiki that either lab's capability has been measured by a third party rather than announced by itself
New in Wiki
Created today. Flagged for your review.
- Military and Intelligence Capability Evals (new — concept page; the first on this wiki for capability measurement in targeting and weapons, deliberately separate from AI-Enabled Cyberattacks because that page covers the cyber surface and this one covers the physical)
- Beyond Solver Verdicts: Generative Reward Models for Autoformalization (new)
- Memory as Plans: World-Action Modeling with Memory-Grounded Planning (new)
Updates
- Anthropic: Frontier Red Team capability evals added to Recent Activity at day +4, with the unresolved 47.2 km attribution carried rather than picked
- AI-Enabled Cyberattacks: cross-linked to the new page as the physical-world half of the same risk surface
- Open-Weights Policy Fight: cross-linked — the first measurement on this wiki of what open weights actually hand over in these tasks
- AI for Mathematics: GenV added to Key Papers as the verification-side counterpart to yesterday's declaration
- Embodied Agents: MaP-WAM added to Related Concepts