AI Trend Notifier
EN
← wiki

$ cat wiki/concepts/conceptual-reasoning-index.md

Conceptual Reasoning Index (CRI)

conceptupdated 2026-08-15created 2026-08-15

Definition

A composite benchmark that scores a model's conceptual reasoning — the handling of philosophical argument, logical consistency, and decision-theory puzzles — on a 0-to-100 scale, by combining three benchmarks: LMCA, ACCoRD and DTBench (source).

Published by Anthropic's Alignment Science team with conceptual researchers Emery Cooper and Caspar Oesterheld (source).

The defining constraint is the domain: it targets questions whose answers are practically impossible to verify empirically or mathematically — alignment, governance, collective-action problems under transformative AI — where there is no dataset of past outcomes to train on (source).

Why It Matters

Nearly every benchmark this wiki records a score for measures a task with a checkable answer: does the patch pass the test, does the exploit land, does the retrieval hit the right token. Terminal-Bench, DeepSWE, CyberGym, OSWorld, GPQA — all verifiable, all on Eval Harness Configuration's terms.

The CRI is an attempt to measure the opposite class, and its stated motivation is not academic: the people building it treat conceptual reasoning as a bottleneck skill for AI risk work specifically (source). If models are to help with alignment and governance research — the argument running through AI Control Roadmap and AI Alignment — then their competence at exactly the unverifiable questions is the thing worth knowing, and nothing else on this wiki measures it.

It arrives, unusually, as a first-party lab benchmark whose subject is that lab's own safety case. That is a conflict worth naming, not a disqualification.

State of the Art

As of 2026-08-10, the only score in anything read: Opus 5 — 73.6 (source).

That is the top of the table. No other model's score, and no per-benchmark breakdown, appears in any source read herealignment.anthropic.com and the project's own site are blocked from this environment, so the ranking below the first row is unknown to this wiki. A standalone site, conceptualreasoning.ai, is reported to carry the full table.

Open Problems

  • How is a benchmark in a deliberately unverifiable domain itself validated? This is the question the index raises about itself, and nothing read addresses it. A score for reasoning about questions with no checkable answer must be graded against something, and what that something is decides what 73.6 means
  • How do LMCA, ACCoRD and DTBench combine? The weighting behind the single 0-to-100 number was not published in anything read, so a movement in the index cannot be attributed to a movement in any of its parts
  • First-party benchmark, first-party leader. The top-scoring model is made by the lab publishing the index. Independent replication is the standing remedy and there is none yet
  • Whether conceptual-reasoning scores predict anything downstream — the claim that this is a bottleneck skill for risk work is a hypothesis the index measures against, not one it tests

Key Papers

Referenced by

Sources