$ cat wiki/papers/2026/2610.02122-argo-bench.md
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
TL;DR
A 235-table, 7.5-billion-row simulated ERP warehouse where the agent is graded on what its actions do in the simulator, not on whether its SQL matches a key. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of 210 tasks and averages 59.5 points. The design choice that makes it interesting is that the simulator's ground truth is withheld from the warehouse the agent sees — facts have to be reconstructed before they can be acted on.
Authors & Org
unknown. arxiv.org answers EGRESS_BLOCKED from this pipeline, so the paper was
not read and the abstract is the only text available, via the HuggingFace Daily
Papers snapshot of 2026-10-05
(source). No author
list or affiliation appears in anything read and none is guessed.
None of the 14 models is named in anything read, so no figure here attaches to any model page in this wiki.
Method
The stated motivation is a criticism of the incumbent benchmarks, and it is sharper than the usual one: established text-to-SQL benchmarks "evaluate query generation alone, and audits have found their answer keys frequently wrong". Because real warehouses are too sensitive to publish, those benchmarks are built on public datasets "where a business event fits in a single table" — which is not the shape of the problem.
Argo-Bench's answer is to simulate the business instead of borrowing a dataset:
| Component | Scale |
|---|---|
| Tasks | 210 data-science and analytics tasks |
| Simulated world | food-delivery platform, New York City |
| Orders simulated (2024) | 81 million |
| Warehouse tables | 235 |
| Warehouse rows | 7.5 billion |
| Schema modelled on | Oracle E-Business Suite |
| The world is built from *"public data, peer-reviewed industry literature, and | |
| regulatory filings"*, with *"grounded economics, fraud patterns, and marketplace | |
| incentives"*. |
Two design decisions carry the evaluation:
- Ground truth is withheld. The simulator's state is not in the warehouse the agent queries, so "tasks require reconstructing facts by navigating the warehouse before acting on them".
- The agent acts, and the consequences are graded. It files actions — "banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay" — and "the grader scores each by its consequences in the simulator".
Every task carries "an executable reference solution that demonstrates solvability using only the warehouse", which is the claim that a zero is the agent's and not the benchmark's.
Results
All figures from the abstract (source).
| Measure | Value |
|---|---|
| Models evaluated | 14, frontier and open-weight |
| Best model: tasks scoring ≥ 95 | 34.8% of 210 |
| Best model: mean score | 59.5 points |
| The two numbers are the finding, together. A mean of 59.5 with only a third of | |
| tasks near-solved describes an agent that gets partway through most of the work and | |
| finishes about one task in three. A benchmark reporting only the mean would read as | |
| competent. |
Significance
It is the first entry on Agents (LLM Agents) that grades an agent on state it changed rather than on text it produced. That page has accumulated the opposite case repeatedly — most directly MCP — Model Context Protocol's 2026-08-26 entry, where One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (arXiv:2608.19741) scores on "the terminal backend state" a session leaves behind and finds that a well-formed tool call is not evidence the task completed. Argo-Bench applies that principle to analytics, where the conventional measure is the query string.
For Eval Harness Configuration the relevant line is not a score at all. "Audits have found their answer keys frequently wrong" is a claim that an entire benchmark family's ground truth is unreliable — the same class of finding as SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents, which that page already holds for SWE-Bench Pro. Which audits, and which benchmarks, is not stated in anything read, so nothing in this wiki is corrected on the strength of it; what changes is how much a text-to-SQL cell is worth.
The withheld-ground-truth design is the part worth watching. It makes the task two-stage — reconstruct, then act — and a benchmark that cannot be solved by pattern-matching the question to a query is a different instrument from one that can. Whether 59.5 is mostly a reconstruction failure or mostly an action failure is the question the headline figure hides.
Open Questions
- No model is named. 14 models, one aggregate, and nothing attributable — so no model page in this wiki gains a row from this.
- The split between reconstruction and action is not reported, and it is the whole diagnostic value of a two-stage design.
- The grader is the simulator. Nothing read says how action consequences are scored, how partial credit is assigned, or how a 59.5 decomposes.
- Harness unspecified. Per Eval Harness Configuration, a 34.8% on an agentic benchmark without the scaffold, turn cap and tool surface stated is not yet a comparable figure, and none is given.
- Public availability is not stated — no release, licence or access route for the warehouse or the 210 tasks appears in anything read.
Cite
arXiv:2610.02122, published 2026-10-01. Surfaced via HuggingFace Daily Papers, 2026-10-05, 26 upvotes — a popularity signal from that community and nothing more (source).