$ cat wiki/papers/2026/2610.07557-checkerbench.md
CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
TL;DR
CheckerBench tests whether an agent can build a working static-analysis checker across a whole repository workflow, rather than only identify or patch a defect. (source)
Authors & Org
Hang He, Li Wang, Hao Chen, Yuchen Shao, Yuling Shi, Lisheng Wang, Peiyang Liu, Goose Lin, Zaiyuan Wang, Haiying Sun, Ting Su and Chengcheng Wan. Affiliations are not stated in the captured abstract record. Submitted 2026-10-06. (source)
Method
The benchmark contains 300 tasks from 297 CVEs, 167 repositories and 85 CWEs across five language ecosystems. Each includes vulnerable and fixed revisions, a pinned analysis environment and a checker scaffold. CheckerLab independently rebuilds submissions and evaluates diagnostic contrast between the revisions, patch localization, false positives and tool use. (source)
Results
Across 21 model-harness configurations with three independent repeats each, mean Pass@1 is 32.30%; the best configuration reaches 45.33%. The abstract does not name that configuration. (source)
Significance
For Agents (LLM Agents), this extends evaluation to reusable checking tools. Read alongside Eval Harness Configuration: the reported unit is a model-harness configuration, so the results do not isolate model quality. (source)
Open Questions
Which configurations lead, and how much do compilation feedback, false-positive tolerance and tool budgets explain their differences? The abstract cannot answer those questions; the full paper was not read. (source)
Cite
Hang He et al. (2026). CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers? arXiv:2610.07557.