$ cat wiki/papers/2026/2610.05622-undobench.md
UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
TL;DR
UndoBench separates nominal task completion from recovery after operational faults. In its frozen lost-acknowledgment study, nominal competence reaches 83.54%, while conditional recovery succeeds at 46.72%. (source)
Authors & Org
Dolly Sah, Tanmay Sah, Harshul Jain and Tanya Sah. The abstract page does not establish affiliations. Submitted 2026-10-04; read scope is the abstract and bibliographic record, not the full paper. (source)
Method
The benchmark spans 36 base workflows, 36 fault scenarios and 8 enterprise domains. Counterfactual paired trials use identical seeds, with effect-history and environment-state oracles to distinguish what an agent attempted from what actually changed. (source)
The frozen study covers 12 held-out workflows, two open-weight models, two frameworks and three recovery paradigms, totaling 5,760 executions / 2,880 paired trials. (source)
Results
Naive retry produces duplicate external effects in 53.33% of trials. Recovery depends on the execution boundary: before mutation, methods behave similarly without duplicates among capable trials; during partial mutation, the evaluated strategies fail on composite workflows; after commit but before acknowledgment, verification and server-side idempotency improve safety. Commercial API extensions reproduce the competence–recovery separation, but the abstract names no model-specific scores. (source)
Significance
For Agents (LLM Agents), this is evidence for evaluating recovery independently of ordinary completion. Interpretation: a successful nominal benchmark run does not establish that retrying a failed tool call is safe. The evidence concerns the evaluated workflows, not every deployed agent. (source)
Open Questions
The abstract does not name the tested models or frameworks, quantify the commercial extensions, or provide uncertainty intervals. Those details require reading the full paper before comparing model vendors. (source)
Cite
Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah. “UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents.” arXiv:2610.05622, 2026. arXiv.