AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2610.12289-testprism.md

TestPrism: Rethinking Test Evaluation Beyond a Single Reference

paperupdated 2026-10-10created 2026-10-10

TL;DR

Tests can accept one reference solution while rejecting other valid implementations. TestPrism evaluates both acceptance of valid programs and rejection of invalid ones, exposing a gap hidden by single-reference scoring. (source)

Authors & Org

Han Li, Lingxiang Hu, Jiacheng Huang, Ziqian Jiang, Jingkai Luo, Wei Gao, Yunfan Tan, Zun Wang and Jiaheng Liu. Affiliations are not established by the captured abstract record. Submitted 2026-10-08. (source)

Method

The benchmark contains 300 tasks from 17 sources and 3000 candidate implementations, evenly divided between valid and invalid solutions. Joint Success Function requires tests to fail on the initial program, accept every valid candidate and reject every invalid candidate. TestHelix combines heterogeneous test/repair synthesis, peer cross-validation and recursive self-improvement. (source)

Results

Across fourteen baseline coding-agent configurations, the abstract reports 28.00% Joint Success Function versus 59.67% single-reference success. TestHelix improves the joint metric by 8.67 to 9.00 percentage points over native harness comparators across two models. The captured abstract does not name those models or give evaluation cost. (source)

Significance

For Agents (LLM Agents) and Eval Harness Configuration, the practical implication is that test validity requires more than detecting the original bug: a test must also tolerate legitimate alternative implementations. This is an interpretation of the benchmark design, not independent replication. (source)

Open Questions

  • How representative are the valid and invalid candidate implementations of production code? (source)
  • Does TestHelix retain its gain when inference cost is matched? The abstract does not answer this. (source)

Cite

arXiv:2610.12289. Evidence scope: title, authors and abstract; the full paper was not read. (source)

Referenced by

Sources