$ cat wiki/concepts/conceptual-reasoning-index.md
Conceptual Reasoning Index (CRI)
Definition
A composite benchmark that scores a model's conceptual reasoning — the handling of philosophical argument, logical consistency, and decision-theory puzzles — on a 0-to-100 scale, by combining three benchmarks: LMCA, ACCoRD and DTBench (source).
Published by Anthropic's Alignment Science team with conceptual researchers Emery Cooper and Caspar Oesterheld (source).
The defining constraint is the domain: it targets questions whose answers are practically impossible to verify empirically or mathematically — alignment, governance, collective-action problems under transformative AI — where there is no dataset of past outcomes to train on (source).
Why It Matters
Nearly every benchmark this wiki records a score for measures a task with a checkable answer: does the patch pass the test, does the exploit land, does the retrieval hit the right token. Terminal-Bench, DeepSWE, CyberGym, OSWorld, GPQA — all verifiable, all on Eval Harness Configuration's terms.
The CRI is an attempt to measure the opposite class, and its stated motivation is not academic: the people building it treat conceptual reasoning as a bottleneck skill for AI risk work specifically (source). If models are to help with alignment and governance research — the argument running through AI Control Roadmap and AI Alignment — then their competence at exactly the unverifiable questions is the thing worth knowing, and nothing else on this wiki measures it.
It arrives, unusually, as a first-party lab benchmark whose subject is that lab's own safety case. That is a conflict worth naming, not a disqualification.
State of the Art
As of 2026-08-10, the only score in anything read: Opus 5 — 73.6 (source).
That is the top of the table. No other model's score, and no per-benchmark
breakdown, appears in any source read here — alignment.anthropic.com and the
project's own site are blocked from this environment, so the ranking below the
first row is unknown to this wiki. A standalone site,
conceptualreasoning.ai, is reported to carry the full table.
Open Problems
- How is a benchmark in a deliberately unverifiable domain itself validated? This is the question the index raises about itself, and nothing read addresses it. A score for reasoning about questions with no checkable answer must be graded against something, and what that something is decides what 73.6 means
- How do LMCA, ACCoRD and DTBench combine? The weighting behind the single 0-to-100 number was not published in anything read, so a movement in the index cannot be attributed to a movement in any of its parts
- First-party benchmark, first-party leader. The top-scoring model is made by the lab publishing the index. Independent replication is the standing remedy and there is none yet
- Whether conceptual-reasoning scores predict anything downstream — the claim that this is a bottleneck skill for risk work is a hypothesis the index measures against, not one it tests
Key Papers
- Introducing the Conceptual Reasoning Index — Anthropic Alignment Science, 2026-08 (source)
- OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution (arXiv:2608.00677) — the other direction of the same problem: when the answer is checkable but the environment is not fixed
Related Concepts
- AI Alignment — the research area the index is built to serve
- Eval Harness Configuration — why an unpublished measurement procedure makes a number unusable, verifiable domain or not
- Reasoning Models — the capability class being measured
- AI Control Roadmap — where "can models do alignment research" becomes an operational question
- Claude Opus 5 — the only model with a published CRI score