AI Trend Notifier
EN
← wiki

$ cat wiki/concepts/military-intelligence-evals.md

Military and Intelligence Capability Evals

Definition

Evaluations that measure how well a model performs the specific technical tasks of intelligence targeting and conventional weapons development — locating a person from fragmentary information, correlating identities across platforms, guiding a drone onto a moving vehicle — rather than measuring whether a model will agree to help with them.

The distinction is the whole point and it is not a fine one. A refusal test asks what the model is willing to do; a capability eval asks what it can do, and therefore what an open-weights release, a jailbreak or a stolen API key actually hands over. The two answers come apart: a model that reliably refuses can still be, in capability terms, an expert.

This is a separate risk domain from cyber. AI-Enabled Cyberattacks covers intrusion, malware and influence operations; this page covers the physical-world targeting and weapons-engineering tasks that Anthropic's Frontier Red Team describes as previously understudied relative to the longstanding focus on cyber, biological and nuclear risk (1 pass) (source).

Why It Matters

A capability number is the only part of this argument that survives a weight release. Every other safeguard a lab has — refusal training, classifiers, account termination, the enforcement recorded on AI-Enabled Cyberattacks — is attached to the lab's platform. Once weights are public, none of it travels with the model; the capability does. So a measurement of what the weights can do is the only thing that tells an open-weights decision what it is deciding, which makes this the missing input to Open-Weights Policy Fight.

It also makes "behind the frontier" a checkable claim rather than a reassurance. Open-weights models from PRC developers tested in this work were found "behind the frontier, but also showed concerning ability to identify and target adversaries, and improve weapon performance" (2 passes, near-verbatim both times). Behind the frontier and sufficient for the task are not the same finding, and only the second one matters to someone deciding whether a release is safe.

And it changes what a safety case has to contain. This wiki's Preparedness Framework pages record thresholds defined for cyber, bio and nuclear. A fourth and fifth domain with published numbers is a fourth and fifth place a threshold has to be set — or visibly not set.

State of the Art (2026-09-10)

Anthropic's Frontier Red Team publishes evals in two domains, and the geolocation number is the one that moved

Measuring AI capabilities in intelligence targeting and conventional weapons, 2026-09-10, from Anthropic's Frontier Red Team. Captured here 2026-09-14, day +4www.anthropic.com and every secondary outlet attempted answer EGRESS_BLOCKED, so nothing below was read first-party and each claim carries a pass count (source).

Targeting — three evals, the same list in 2 passes: identity correlation on synthetic multi-platform social data, photo geolocation against human GeoGuessr baselines, and text geolocation with a sandboxed search tool.

The photo geolocation result is the one with numbers behind it. Across a stated 6,000 photos (1 pass), Mythos Preview reports a median distance error of 37.0 km (2 passes agree on 37 for this model), against Champion Division GeoGuessr players — stated as the top 0.01% of the player base — at 151 km (1 pass). A second model reports 47.2 km, and which model that is was not resolved: one pass attributes it to Opus 5, another to Mythos 5, and both are recorded in ## Conflicting Reports below rather than picked between.

Why 151 km is the number to read against: the human baseline is not a casual player, it is the top 0.01%, and the model's median error is roughly a quarter of it. Coverage's own gloss — "approaching superhuman capabilities for geolocating outdoor photos" (1 pass) — is the write-up's, not this wiki's, but the ratio does not depend on the gloss.

Conventional weapons — three tasks, in simulation (1 pass, explicit): terminal guidance to a vehicle, payload drop within a grenade-like radius, and GPS-denied / spoofed navigation. Two metrics: the share of flights whose median miss lands within five metres, described as "the approximate lethal radius of a grenade", and the median miss distance (1 pass for both definitions).

The model ordering is reported and the table behind it is not. Named as separating on these tasks: Claude Opus 5, Mythos-class (Claude Mythos Preview), Claude Sonnet 5, and open-weights Kimi K3 and GLM-5.2 (1 pass). Kimi K3 is the only non-Anthropic model any pass gives figures for: a 83% hit rate with a 0.4 m median miss — a lower hit rate than Sonnet but a better median miss (1 pass). No figure for any other model was recovered in any pass.

That split is worth keeping even at one pass, because it is the shape a single-number summary would destroy: a model can be less reliable at the task and more precise when it succeeds, and "behind the frontier" collapses both into one word.

The mitigation is named and unmeasured

"On-platform safety measures are necessary, like new classifiers that have been implemented to block such misuse" (2 passes), described as expanding guardrails "beyond the longstanding focus on cyber, biological, and nuclear risks" (1 pass). No classifier name, deployment date, precision or recall figure, or coverage statement appeared in any pass. This is the same auditability gap AI-Enabled Cyberattacks records for Google's Fairwind gate on 2026-09-03: the remedy is announced, and there is nothing published that anyone outside the lab could check it against.

And it is on-platform by construction, which is exactly the boundary the open-weights half of the same post crosses. The post measures a risk that survives weight release and answers it with a control that does not.

The same-day threat report is a different document

Detecting and countering misuse of AI: September 2026 was published 2026-09-10 as well and is held separately (source); its findings live on AI-Enabled Cyberattacks. Coverage repeatedly conflates the two. The evaluation post states that "models have become useful to actors seeking to misuse our platform for surveillance and conventional weapons development" (2 passes), and write-ups dated 2026-09-10 and 2026-09-12 attach that to the Russian self-targeting drone operation from the threat report — but no pass quotes the evaluation post naming that incident, so the connection is recorded as the coverage's and not as the post's.

Open Problems

  • Which model holds 47.2 km. One pass each for Opus 5 and Mythos 5; nothing read settles it. Recorded below.
  • No first-party text has been read at all. Every figure on this page is a search extract. The eval names, the 6,000-photo sample, the five-metre radius and every number should be re-read against the post itself the first time anthropic.com is reachable from a run.
  • Are the evals reproducible by anyone else? No pass states that a harness, prompt set or scoring rubric is published. A capability claim nobody outside the lab can re-run is a claim about the lab's own measurement, and the open-weights finding in particular is one other people have standing to check.
  • No threshold is attached to any number. 37.0 km is a measurement, not a verdict; nothing read says what value would trigger what response under Preparedness Framework.
  • Only two open-weights models were named. Both Chinese. Nothing read states whether Western open-weights models were tested and omitted, or not tested.
  • Does a classifier survive the thing it is defending against? The mitigation is on-platform; the finding is about weights that leave the platform.

Key Papers

  • Measuring AI capabilities in intelligence targeting and conventional weaponsAnthropic Frontier Red Team, 2026-09-10 (source) (Anthropic)
  • Detecting and countering misuse of AI: September 2026 — same day, different document; the misuse record rather than the capability measurement (source)

Conflicting Reports

  • The 47.2 km median geolocation error is attributed to two different models. One search pass reads "Mythos Preview and Opus 5 models demonstrated near-superhuman performance in geolocating images, with median errors as low as 37 kilometers"; another reads "Mythos Preview and Mythos 5 achieved median distance errors of 37.0 km and 47.2 km respectively across 6,000 photos". 37.0 km for Mythos Preview is common to both and is the figure this page carries. The second figure's owner — Claude Opus 5 or Mythos 5 — is one pass each and unresolved (source).
  • "First public systematic evaluation framework" for these domains appears in one pass only. Recorded as the coverage's wording and not adopted — a priority claim carried by a single secondary source is the kind this wiki has been wrong about before.

Referenced by

Sources