$ cat wiki/papers/2026/2608.17271-asi-bench.md
ASI-Bench: At the Dawn of Artificial Superintelligence (arXiv:2608.17271)
TL;DR
60 project-level research tasks across 11 scientific domains, built by 40+ experts over 31,000+ human hours, with one design choice that makes it worth reading: it progressively withdraws human methodological guidance within the same project. Across 18 agent–model configurations the average score falls 50.91 → 29.10 → 26.62 as guidance is removed (source).
Authors & Org
Not obtainable. arxiv.org is EGRESS_BLOCKED from this environment and the
paper was not read; the HuggingFace Daily Papers snapshot carries title, id, date
and abstract only
(source).
Listed on HuggingFace Daily Papers, 2026-08-20, 51 upvotes — that community's
popularity signal and nothing more. A submission portal is named in the abstract
(https://asibench.apexin.ai/submit), which is the only affiliation signal read
(source).
Method
| Component | Figure |
|---|---|
| Tasks | 60, project-level |
| Domains | 11 scientific |
| Build cost | 40+ experts, 31,000+ human hours |
| Configurations evaluated | 18 agent–model combinations |
| Validation | expert review, AI-assisted auditing, sandbox execution, scorer validation |
| The distinguishing mechanism is the guidance ladder — three settings applied | |
| to the same research project: |
| Setting | Average score |
|---|---|
| Full methodological guidance | 50.91 |
| Method specified only | 29.10 |
| Agent must determine the method itself | 26.62 |
Results
The score roughly halves the moment full guidance is withdrawn — a 21.81 point drop from setting 1 to setting 2 — and then moves only 2.48 further when method selection is handed to the agent entirely. The paper reads this as evidence that current systems remain heavily dependent on human guidance and are far from autonomously conducting end-to-end, project-level research.
The shape of the curve is more informative than its endpoint, and the paper does not dwell on it: almost the entire loss occurs at the first withdrawal. Whatever the guidance was supplying, it was not primarily which method to pick — if it were, the second step would cost more than 2.48 points.
What the abstract does not give: the 18 configurations, the 11 domains, the score's scale or units, any per-domain breakdown, inter-rater agreement for the expert review, or what the AI-assisted auditing contributed.
Significance
It is the second independent measurement in two days of the same deficit, on a different instrument. How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905) (2026-08-19) ran 8 harness-model combinations over 100 research tasks and located the failure at the model level rather than in any scaffold, naming the missing faculty a metacognitive loop. ASI-Bench varies a different axis — how much human methodology is present, holding the project fixed — and finds the collapse arrives with the removal of human framing, not with the difficulty of the science. Neither cites the other; the pairing is this wiki's and is labelled as such.
Read together they say something narrower and more useful than either alone: agents execute research well when a human has already decided what the research is. That is the same boundary Eval Harness Configuration now records between execution reliability and self-assessment, reached from the autonomy side.
The name is doing work the results do not. A benchmark titled At the Dawn of Artificial Superintelligence reports its subject scoring 26.62 unaided, and the framing — "the first benchmark to jointly evaluate innovative exploration and autonomous scientific execution" — is a claim about novelty that How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905), published four days earlier over 100 tasks and 800 trajectories, complicates. This wiki records the measurements and not the framing.
Open Questions
- What is the score out of? Every figure here — 50.91, 29.10, 26.62 — is unitless in everything read, so the 21.81-point drop cannot be converted into a success rate or compared with any other benchmark on this wiki.
- Which 18 configurations? The "far from autonomous" conclusion depends entirely on the strongest system tested being genuinely strong, and none is named. This is the identical gap recorded on How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905).
- Is a project-level task scoreable at all? Expert review plus AI-assisted auditing plus scorer validation is three instruments; no agreement figure between them was published.
- Open submissions and benchmark integrity — the abstract invites the world to contribute new tasks. A benchmark that grows by public submission after its headline numbers are published cannot keep comparing later systems to those numbers, and nothing read addresses versioning.
- Author list, affiliation, licence — unknown; the paper was not read.
Cite
ASI-Bench: At the Dawn of Artificial Superintelligence (2026). arXiv:2608.17271.