AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2610.07557-checkerbench.md

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

paperupdated 2026-10-08created 2026-10-08

TL;DR

CheckerBench tests whether an agent can build a working static-analysis checker across a whole repository workflow, rather than only identify or patch a defect. (source)

Authors & Org

Hang He, Li Wang, Hao Chen, Yuchen Shao, Yuling Shi, Lisheng Wang, Peiyang Liu, Goose Lin, Zaiyuan Wang, Haiying Sun, Ting Su and Chengcheng Wan. Affiliations are not stated in the captured abstract record. Submitted 2026-10-06. (source)

Method

The benchmark contains 300 tasks from 297 CVEs, 167 repositories and 85 CWEs across five language ecosystems. Each includes vulnerable and fixed revisions, a pinned analysis environment and a checker scaffold. CheckerLab independently rebuilds submissions and evaluates diagnostic contrast between the revisions, patch localization, false positives and tool use. (source)

Results

Across 21 model-harness configurations with three independent repeats each, mean Pass@1 is 32.30%; the best configuration reaches 45.33%. The abstract does not name that configuration. (source)

Significance

For Agents (LLM Agents), this extends evaluation to reusable checking tools. Read alongside Eval Harness Configuration: the reported unit is a model-harness configuration, so the results do not isolate model quality. (source)

Open Questions

Which configurations lead, and how much do compilation feedback, false-positive tolerance and tool budgets explain their differences? The abstract cannot answer those questions; the full paper was not read. (source)

Cite

Hang He et al. (2026). CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers? arXiv:2610.07557.

Referenced by

Sources