$ cat wiki/concepts/rd-automation-index.md
R&D Automation Index
Definition
A measurement instrument that reports what fraction of a frontier lab's own AI research and development is performed by AI, grading each unit of work on an autonomy scale rather than counting outputs. The first and so far only instance is Anthropic's, published 2026-09-17 as a prototype by the Anthropic Institute, alongside two other measurements — oversight of autonomous AI agents and compute allocation (source).
The scale runs AL0–AL5, which one source states is borrowed from Epoch AI (source):
| Level | Definition as published |
|---|---|
| AL0 | No AI involvement |
| AL3 | AI collaborates — large chunks of work under close human direction |
| AL4 | AI leads — completes most of a task end-to-end from a high-level prompt, human supervises |
| AL5 | Fully autonomous, no human in the loop |
| AL1 and AL2 are not described in anything read, so the scale this wiki can state is four of its six points (source). |
Why It Matters
It converts a rhetorical question into a reported number, and the number is about the lab reporting it. Frontier Pacing has, since 2026-07-28, collected commitments and arguments about how fast labs should move; every one of them is a statement of intent. An index that says 26% of this lab's AI R&D is AL4 as of August 2026 is the first quantity on that page that can be restated next quarter and compared (source).
The measurement and the thing measured share an author, and the measurement is itself performed by the model. Anthropic chose the task taxonomy, applied the scale, and used Claude to read the internal records and build the category hierarchy. That is not a reason to discard the figure; it is the reason the figure needs the methodology published beside it, which it was (source).
It is the only published artefact that makes Frontier Pacing step 1 checkable. Amodei's 2026-09-12 plan committed Anthropic to embedded third-party evaluators with permanent, employee-level access and the right to publish; an evaluator with badge access and no agreed instrument measures nothing in particular. This index is a candidate instrument, published five days later, and one source explicitly pairs the two (source) (source).
State of the Art (2026-09-19)
One lab, one prototype, one reading. Anthropic's published figures, as of August 2026 (source):
| Measure | Value |
|---|---|
| AI R&D work at AL4 ("leads") | 26% |
| Measured tasks at AL3 or higher | above 90% |
| Measured work at AL5 (fully autonomous) | none |
| The methodology, as published (source): for each week in July, a 20% sample of employees in departments involved in model R&D; Claude read internal records — Slack messages, documentation — to identify the work those employees performed; that produced roughly 15,000 individual tasks; Claude organised them into a hierarchy of 542 categories. |
Two things in that paragraph do not line up and this wiki does not reconcile them. The sampling window is July and the headline is stated as of August 2026; nothing read explains the gap. And the trajectory is reported two ways — three sources say Claude led less than 1% in February, one headline says 1% in March (source). Both are coverage; the document that would settle either was not read, www.anthropic.com answering EGRESS_BLOCKED.
No second lab has published anything comparable. Anthropic publishes the methodology so that others can measure the same things; as of this page's creation, nothing read reports another lab doing so.
Open Problems
- Who grades. The published criticism is not that the number is wrong but that "who judged which tasks sit at which level, on what evidence" is entirely internal, with no external verification (source). An index whose denominator its author selects is a self-report with arithmetic attached until somebody else runs it.
- What "leads" licenses. One outlet's headline is that "'lead' doesn't mean what you think" (source). AL4 as defined keeps a human supervising every task; 26% at AL4 and 0% at AL5 is a narrower claim than "AI is building AI", which is how several outlets titled it.
- Denominator drift. Nothing read states whether the 542 categories are fixed. If the taxonomy is rebuilt each round by the model, a rising percentage and a shifting definition of "AI R&D work" are indistinguishable from outside.
- The other two measurements have no numbers here. Oversight of autonomous agents and compute allocation are named in every pass and quantified in none. This page holds one of three.
- It measures Anthropic's R&D, not the frontier. A single lab's internal automation rate is evidence about that lab. The stated purpose — minimising the gap between what labs know and what the public knows — needs more than one participant, and has one.
Key Papers
- No paper. The artefact is a lab publication, not a preprint, and no independent replication or critique with a method exists in anything read.
- Adjacent on this wiki: Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement formalises the boundary between a system that improves itself and one improved from outside — the same distinction AL4 and AL5 draw operationally, arrived at from theory rather than from payroll records. Neither cites the other; the adjacency is this wiki's reading.
Related Concepts
- Frontier Pacing — the argument this index is built to supply evidence for
- AI Governance — where a published self-measurement becomes a governance instrument or fails to
- Agents (LLM Agents) — AL4 describes an agent completing a task end-to-end from a high-level prompt, which is that page's subject measured in headcount
- Anthropic — the lab that published it and the lab it measures