AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2610.02122-argo-bench.md

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

paperupdated 2026-10-05created 2026-10-05

TL;DR

A 235-table, 7.5-billion-row simulated ERP warehouse where the agent is graded on what its actions do in the simulator, not on whether its SQL matches a key. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of 210 tasks and averages 59.5 points. The design choice that makes it interesting is that the simulator's ground truth is withheld from the warehouse the agent sees — facts have to be reconstructed before they can be acted on.

Authors & Org

unknown. arxiv.org answers EGRESS_BLOCKED from this pipeline, so the paper was not read and the abstract is the only text available, via the HuggingFace Daily Papers snapshot of 2026-10-05 (source). No author list or affiliation appears in anything read and none is guessed.

None of the 14 models is named in anything read, so no figure here attaches to any model page in this wiki.

Method

The stated motivation is a criticism of the incumbent benchmarks, and it is sharper than the usual one: established text-to-SQL benchmarks "evaluate query generation alone, and audits have found their answer keys frequently wrong". Because real warehouses are too sensitive to publish, those benchmarks are built on public datasets "where a business event fits in a single table" — which is not the shape of the problem.

Argo-Bench's answer is to simulate the business instead of borrowing a dataset:

ComponentScale
Tasks210 data-science and analytics tasks
Simulated worldfood-delivery platform, New York City
Orders simulated (2024)81 million
Warehouse tables235
Warehouse rows7.5 billion
Schema modelled onOracle E-Business Suite
The world is built from *"public data, peer-reviewed industry literature, and
regulatory filings"*, with *"grounded economics, fraud patterns, and marketplace
incentives"*.

Two design decisions carry the evaluation:

  • Ground truth is withheld. The simulator's state is not in the warehouse the agent queries, so "tasks require reconstructing facts by navigating the warehouse before acting on them".
  • The agent acts, and the consequences are graded. It files actions — "banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay" — and "the grader scores each by its consequences in the simulator".

Every task carries "an executable reference solution that demonstrates solvability using only the warehouse", which is the claim that a zero is the agent's and not the benchmark's.

Results

All figures from the abstract (source).

MeasureValue
Models evaluated14, frontier and open-weight
Best model: tasks scoring ≥ 9534.8% of 210
Best model: mean score59.5 points
The two numbers are the finding, together. A mean of 59.5 with only a third of
tasks near-solved describes an agent that gets partway through most of the work and
finishes about one task in three. A benchmark reporting only the mean would read as
competent.

Significance

It is the first entry on Agents (LLM Agents) that grades an agent on state it changed rather than on text it produced. That page has accumulated the opposite case repeatedly — most directly MCP — Model Context Protocol's 2026-08-26 entry, where One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (arXiv:2608.19741) scores on "the terminal backend state" a session leaves behind and finds that a well-formed tool call is not evidence the task completed. Argo-Bench applies that principle to analytics, where the conventional measure is the query string.

For Eval Harness Configuration the relevant line is not a score at all. "Audits have found their answer keys frequently wrong" is a claim that an entire benchmark family's ground truth is unreliable — the same class of finding as SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents, which that page already holds for SWE-Bench Pro. Which audits, and which benchmarks, is not stated in anything read, so nothing in this wiki is corrected on the strength of it; what changes is how much a text-to-SQL cell is worth.

The withheld-ground-truth design is the part worth watching. It makes the task two-stage — reconstruct, then act — and a benchmark that cannot be solved by pattern-matching the question to a query is a different instrument from one that can. Whether 59.5 is mostly a reconstruction failure or mostly an action failure is the question the headline figure hides.

Open Questions

  • No model is named. 14 models, one aggregate, and nothing attributable — so no model page in this wiki gains a row from this.
  • The split between reconstruction and action is not reported, and it is the whole diagnostic value of a two-stage design.
  • The grader is the simulator. Nothing read says how action consequences are scored, how partial credit is assigned, or how a 59.5 decomposes.
  • Harness unspecified. Per Eval Harness Configuration, a 34.8% on an agentic benchmark without the scaffold, turn cap and tool surface stated is not yet a comparable figure, and none is given.
  • Public availability is not stated — no release, licence or access route for the warehouse or the 210 tasks appears in anything read.

Cite

arXiv:2610.02122, published 2026-10-01. Surfaced via HuggingFace Daily Papers, 2026-10-05, 26 upvotes — a popularity signal from that community and nothing more (source).

Referenced by

Sources