$ cat wiki/papers/2026/2610.07753-safeactbench.md
From Evidence to Action: How Tool-Using Agents Fail
TL;DR
SafeActBench separates a correct final outcome from having established the evidence needed before acting. Strong static judgment can coexist with weak interactive execution. (source)
Authors & Org
Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai, Mong-Li Lee and Wynne Hsu. Affiliations are not established by the captured abstract record. Submitted 2026-10-06. (source)
Method
The benchmark contains 656 cases across six operational domains and five protocols, spanning static action judgment, investigated non-action, single actions and dependent workflows. Its provenance-bound Evidence Ledger and deterministic evaluator track what evidence was established, when actions occurred and whether prerequisites were met. (source)
Results
Across ten model-harness configurations, failures include incomplete investigation and premature action. Single actions are usually reliable after obtaining the required evidence, while workflows also expose unresolved dependencies and incomplete execution. The abstract provides no numerical success rate or model ranking. (source)
Significance
For Agents (LLM Agents), the evaluation target includes the order of evidence and actions. Compared with UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents, this asks whether action was justified beforehand; recovery after an external effect is a separate question. (source) (recovery source)
Open Questions
How do the tested models differ, and how well do the six domains transfer to deployment? The full paper and project implementation were not read, so the abstract's qualitative finding is not a deployment guarantee. (source)
Cite
Hongzhan Lin et al. (2026). From Evidence to Action: How Tool-Using Agents Fail. arXiv:2610.07753.