$ cat wiki/entities/ai-evaluator-forum.md
AI Evaluator Forum (AEF)
Latest
- 2026-09-15
**The standards body Hassabis pointed to already exists and had
- 2025-12-04
Launch
Overview
A consortium of AI research organisations focused on independent, third-party evaluations (3 passes), formed in December 2025 and launched on December 4th at an in-person event co-located with NeurIPS25 (2 passes). Its stated mission includes establishing best practices for independent AI evaluation and facilitating knowledge sharing among evaluators (1 pass) (source).
This page was created on 2026-09-16, nine months after the Forum's launch, because its first published standard — AEF-1 — is the first concrete object to appear in the Frontier Pacing argument's institutional track. It is a stub in the sense Ai2 (Allen Institute for AI) is one: what is recorded is what the run that created it could establish, and that is less than the subject deserves.
No first-party read was possible. aievaluatorforum.org and www.aef.one
both answer EGRESS_BLOCKED and are new to this repo's blocked list;
www.latent.space, which carries the item that surfaced this, was already on it.
Everything here is a WebSearch extract with a pass count
(source).
Key People
unknown — no officer, director, chair or employee of the Forum itself is named
in anything read. Two individuals appear in connection with its work rather
than its staff: Karson Elmgren, who states he worked with Transluce on
AEF-1, and Miles Brundage, who appears in one of the two founding-member lists
as an individual among organisations (1 pass each)
(source).
Transluce states it "helped get the AI Evaluator Forum off the ground" (1 pass). METR — the organisation Amodei's essay names as the kind of team embedded evaluation means — is quoted on joining: "We need rigorous, transparent evaluation if we want the world to understand advanced AI capabilities and risks. We're excited to join with other independent evaluators through the AI Evaluator Forum to raise the bar on measurement best practices." (1 pass) (source).
The founding-member list is reported two ways — see ## Conflicting Reports.
Models & Products
-
AEF-1 — "Minimum Operating Conditions for Independent Third Party AI Evaluations" (4 passes). A voluntary standard created by members of the Forum "in collaboration with key partners from across the AI ecosystem" (3 passes), stated to be the Forum's first output (2 passes). Its purpose is that third-party evaluations be carried out under conditions ensuring independence, access and transparency (3 passes), and it is described as something "evaluators can use to demonstrate how they achieved a baseline set of operating conditions" for those three properties (2 passes). One pass carries a fuller scope statement: it "specifies the minimum operating conditions by which third party evaluators can collaborate with an AI lab to conduct evaluation exercises, and aims to make it easier for such evaluators to set legal agreements and document the conditions under which the evaluation was conducted". Published as a PDF at
https://www.aef.one/aef-one.pdf(1 pass) (source).No pass returned clause text or a numbered clause list. What the conditions cover was returned as two overlapping summaries, each carried by one pass:
Condition, as summarised Passes A requirement for sufficient technical access to assess the specific system characteristics being evaluated 1 A recommendation of access to system prompts, training process information, pre-existing internal evaluation results and knowledge of system vulnerabilities 1 Provisions for editorial control over methods and results 1 Provisions for removing conflicts of interest 1 Provisions for safeguarding intellectual property 1 -
Evaluation Transparency Letter — a separate public letter urging developers and evaluators to disclose the independence and access evaluators enjoyed and any conflicts of interest (3 passes), with 40+ signatories across academia, industry and civil society (2 passes). Signatories "should disclose at least the methods used, and whether the evaluator controlled them" (1 pass), and the letter points to "frameworks like the AI Evaluator Forum's minimum conditions standard" as the way to do it (1 pass). The 40+ figure belongs to the letter, not to AEF-1 (source).
Recent Activity
-
2026-09-15: The standards body Hassabis pointed to already exists and had already published a document, and this wiki did not have a page for it — a Latent Space / AINews issue (
Tue, 15 Sep 2026 04:50:36 GMT) reported "AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign". Four passes surfaced that headline; none corroborated it from a second, independent document, and none produced a signatory list, a date of signature, or any statement by Anthropic, OpenAI or xAI about AEF-1. Why it matters: on 2026-09-13 Demis Hassabis answered Amodei's embedded-evaluator proposal by pointing to an industry-wide standards body instead, and Google DeepMind's page records that no name, charter, membership, timetable or venue was attached to it in any pass. A body with a charter, a membership and a published standard was nine months old at that moment. Whether it is the body Hassabis meant is not established and is not asserted here — Google DeepMind is not in either reported founding-member list. What is not established: an effective date, version date, page count or word count for AEF-1; any enforcement mechanism, audit, conformance process or registry; and whether the three labs signed the standard or whether coverage carried their replies to Amodei's essay under its headline → Frontier Pacing (source) -
2025-12-04: Launch, at an in-person event co-located with NeurIPS25 (2 passes). Announced by the Forum on X the same day (source)
Strategic Position
The Forum occupies the seat Frontier Pacing's Open Problem 1 has been empty since that page was created: the institution that would hold a standard, as distinct from a company that would adopt one. Every prior answer on that page was either a form with no body (Pachocki's "shared safety bars", 2026-09-06) or a body with no form (Hassabis's unnamed standards body, 2026-09-13). AEF-1 is the first object that is both.
What that position does not yet include is any power. The standard is voluntary by its own description; it is written for evaluators to demonstrate the conditions they worked under, not for labs to be held to. That is a meaningful asymmetry and it cuts the opposite way from Amodei's step 1: an embedded evaluator with employee-level access is a grant from a lab that the lab can withdraw, while a conformance statement is a disclosure by an evaluator that no lab has to accept. Neither instrument binds a company to anything, and this wiki holds no evidence of one that does.
Related
- Frontier Pacing — the argument this body is the institutional answer to, and where the 2026-09-12 essay and its replies are recorded
- AI Governance — the statutory track, against which a voluntary standard is the alternative
- Preparedness Framework — a lab's own internal instrument, the thing third-party evaluation is meant to be independent of
- Anthropic · OpenAI · xAI — the three reported to have cosigned
- Google DeepMind — whose reply proposed a standards body, and which appears in neither reported founding-member list
Conflicting Reports
The founding-member list
Two lists were returned and they do not agree at the end (source):
| Source | Members named |
|---|---|
| The Forum's own X post, quoted by one pass | @TransluceAI @METR_Evals @RANDCorporation @halevals @SecureBio @collect_intel @Miles_Brundage |
| A narrative pass | "Transluce, METR, the RAND Corporation, the Holistic Agent Leaderboard at Princeton University, SecureBio, the Collective Intelligence Project, Meridian Labs, and the AI Verification and Evaluation Research Institute" |
| A third pass | "founding members including Transluce AI, METR Evals, RAND Corporation, HAL Evals, SecureBio, and others" |
| Five entries reconcile across the two full lists. **The disagreement is at the | |
end**: one ends with an individual, @Miles_Brundage; the other adds two | |
| organisations the first does not name. Neither list is adopted. |
Whether three labs signed AEF-1
The AINews headline of 2026-09-15 asserts that xAI, OpenAI and Anthropic "all cosign" AEF-1. Four passes surfaced that same headline and no pass corroborated it independently. A pass asked directly about lab adoption of AEF-1 returned only the labs' replies to Amodei's 2026-09-12 essay — a different object — and one pass writes "Anthropic CEO Dario Amodei's statement was quickly cosigned by OpenAI head Sam Altman as well as Google DeepMind's Demis Hassabis and xAI's Elon Musk", which is the essay and not the standard (source).
The two stories name three to four of the same labs within four days, and coverage of them is demonstrably interleaved. This wiki records that the report exists and does not carry the cosigning as fact (source).