AI Trend Notifier
EN
← wiki

$ cat wiki/concepts/embedded-evaluation.md

Embedded Evaluation

conceptupdated 2026-09-20created 2026-09-20

Definition

Embedded evaluation is third-party safety assessment performed from inside a frontier AI company rather than from outside it. The evaluator is given standing access to the development process — training runs in progress, deployment decisions, staff — instead of being handed a finished model and asked what it can do.

The distinction that defines it is what is measured. External evaluation measures an artefact: here is a checkpoint, here is what it scores. Embedded evaluation measures a process: how the checkpoint came to be, which decisions were taken along the way, and whether the company's stated safety commitments were actually applied. Sam Altman and Dario Amodei both use the phrase "employee-level access" for the access this requires (source).

It is not auditing, and the difference is the absence of a standard. An auditor works to a published standard against which the audited party can be found wanting. As of 2026-09-20 no such standard exists for frontier AI, and Anthropic says so in the announcement of its own first engagement (source). What an embedded evaluator may see, what it must report, to whom, and who pays for it are decided in each bilateral arrangement, by the party being evaluated.

This page was created 2026-09-20, eight days after the mechanism was named and two days after the first contract for one was signed. It exists separately from Frontier Pacing — which holds the argument that pacing is necessary — because the mechanism now has a counterparty, a price and a published set of objections, and those are facts about the instrument rather than about the case for it.

Why It Matters

Every safety claim this wiki holds about a frontier model is the developer's own claim about its own model. Eval Environment Containment records nine incidents across four labs, and the provenance of each is instructive: one surfaced because a third party noticed, three because a competitor disclosed, two because a vendor or evaluator reported in, one because outside researchers scanned the open internet, one because a lab was assembling an evidence package for METR — and the ninth, on 2026-09-18, because the vendor went back through its own logs. Not one was surfaced by the lab's routine monitoring of its own evaluations.

That is the gap embedded evaluation is proposed to close, and it is why access to the process rather than the artefact is the load-bearing part. Every one of those nine incidents was invisible in the finished model.

State of the Art (2026-09-20)

The mechanism has exactly one signed engagement, and it was signed two days ago.

The first contract: Anthropic × Accenture (2026-09-18)

FieldValue
LabAnthropic
EvaluatorAccenture, work led by Faculty, its specialist AI business
Moneyeach expects to invest at least $1 billion over five years
Access"comparable to an employee's"
Scopeevaluating and red-teaming models, alignment assessments, testing safeguards, verifying safety commitments, identifying blind spots, reporting incidents
Who paysAnthropic funds Accenture's work directly
Exclusivitynon-exclusive — more evaluators promised "in the coming weeks"
(source)

One pass reports Accenture acquired Faculty in January 2026. Anthropic also says it is in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding — a materially different funding structure from the Accenture arrangement, and the only place in the announcement where the evaluator is not paid by the evaluated.

This closes three of the four gaps Frontier Pacing recorded against Amodei's step 1 on 2026-09-12 — that commitment had no named counterparty, no contract and no engaged evaluator. It does not close the fourth: no start date appears in anything read.

And it does not carry forward the clause with teeth. Amodei's essay committed to evaluators holding the right to publish key findings without Anthropic editorial control, subject to narrow security or legal redactions (source). No pass on the Accenture announcement mentions publication rights at all. That is recorded as not-established rather than as a retraction — the essay's commitment stands unless something contradicts it — but it is the single most consequential term of the arrangement and it is missing from the first arrangement made under it.

The same day, a definition of independence the first contract does not meet

More than 100 AI researchers, including Geoffrey Hinton, published a letter on 2026-09-18 stating that embedded evaluators should be "meaningfully independent", and giving three criteria (source). Read against the Accenture arrangement:

Criterion (quoted)Accenture
"should not be owned or governed by frontier AI companies"Met — Accenture is independently owned
"should not have other significant commercial business with them"Not met. One pass reports a joint business group, ~30,000 Accenture staff trained on Claude and tens of thousands using Claude Code (single-sourced figures)
"should not accept any form of payment or other reward contingent on the evaluator's findings"Not established. Anthropic funds the work directly; nothing read says whether any term is contingent on findings
**The two documents are not presented as a dispute and this page does not make
them one.** Nothing read states that the letter names Accenture, that its
authors knew of the partnership, or which was published first. What is
established is that on a single day the mechanism acquired its first instance
and its first written independence criteria, and that **the instance satisfies
one of the three criteria outright**.

The letter was organised by the AI Evaluator Forum (AEF), and names Conrad Stosz as its chair — the first officer of that body named in anything this wiki has read, on a page that carries unknown under Key People. Signatories include representatives of Johns Hopkins, Stanford and METR (1 pass each).

The prior art nobody called by this name

Anthropic's eight-week METR agreement of 2026-09-09 — broad access to transcripts and to staff — is embedded evaluation in everything but the label, and it predates the label by three days (source). It is the narrower instrument (one incident class, a fixed eight weeks, no stated publication terms) and the more independent one (METR is a nonprofit, and in the Accenture announcement it is the party funding itself).

The discovery route matters more than the instrument. That agreement is also how the eighth containment incident was found: Anthropic located the January 2026 transcripts while assembling material to share with METR. An evaluator that had not yet started work had already caused a lab to look somewhere it had not looked.

Open Problems

  • The evaluator is paid by the evaluated, and the announcement says this is wrong. Anthropic funds Accenture directly and, in the same document, concedes the industry has no settled answer on who should fund the process (source). The letter's third criterion addresses only payment contingent on findings, which is a narrower bar than payment as such — and no proposal read here says where the money should come from instead.
  • Publication rights are the whole mechanism and they are unrecorded in the first contract. An evaluator that cannot publish is an internal compliance department with an outside employer, which is precisely the objection one pass raises. Amodei's essay granted the right; the Accenture announcement, as read, does not mention it.
  • Selection is unilateral. The company being evaluated chooses its evaluator, defines the access, and is free to disregard the conclusions. Nothing in any document read constrains any of the three.
  • No standard exists, so no two engagements are comparable and none can be failed. AI Evaluator Forum (AEF)'s AEF-1 is the only published candidate, and this wiki cannot confirm that any lab has signed it — see the ## Conflicting Reports section of Frontier Pacing.
  • Legal protection is absent. The letter states evaluators lack "the independence, resources, and legal protections" needed. No whistleblower protection, safe-harbour, or indemnity for an evaluator appears in anything this wiki holds.
  • Three of four labs have committed to nothing. Altman adopted step 1 by reply on 2026-09-13; Hassabis called for a standards body instead; Musk endorsed the direction; Suleyman conditioned support on evaluators being "truly third-party". Only Anthropic has an engagement, and Altman's commitment has no counterparty (source).

Key Papers

None. This concept has no literature — every document on this page is a company announcement, an essay, or a letter. That is itself worth recording: the mechanism being proposed to verify frontier safety claims has, as of 2026-09-20, no published methodology, no evaluation of its own effectiveness, and no prior case anyone has written up.

Referenced by

Sources