AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2609.33987-opera.md

Opera — persistent feedback for coding agents

paperupdated 2026-10-11created 2026-10-11

TL;DR

Opera keeps a critic’s correction active until the underlying problem is resolved. Its abstract reports gains during inference and after training on critic-guided rollouts. (source)

Authors & Org

Kai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan, Yifan Zhang, Dingjie Song, Dimitris N. Metaxas, Silvio Savarese, Ran Xu and Zeyuan Chen. Affiliations are not stated in the captured abstract page. First submitted 2026-09-27; the reviewed version was revised 2026-10-08. (source)

Method

Periodic and event-driven triggers initiate review. Typed diagnostic operators identify problems, an evidence audit checks feedback before delivery, and persistent notes track subsequent actions to distinguish compliance with advice from resolution of the problem. The authors also use these approximately on-policy rollouts for fine-tuning. (source)

Results

  • Across four policy models, the authors report resolve-rate gains of up to 12.4 percentage points on Terminal-Bench 2.1, 15.0 on a SWE-Bench Pro subset, and 8.9 on DeepSWE v1.1. The abstract also reports gains when a model critiques itself. (source)
  • Fine-tuning Qwen3.5-9B improves held-out SWE-Bench Pro resolve rate by 10.2 percentage points without an inference-time critic. The authors report matching stronger-model rollout training while retaining performance when switching from Openhands to Terminus-2; the comparator degrades under that switch. (source)

Significance

For Agents (LLM Agents), this offers an explicit record of unresolved corrections, complementing the outstanding-requirement record in RunningTab — tracking unfinished workspace requirements. Interpretation: generating useful advice and checking whether it fixed the problem are separate responsibilities. (source) (RunningTab)

Open Questions

The abstract does not provide absolute resolve rates, per-model rows, or a cost-matched comparison. The reported maxima should not be read as gains for every model. Full paper and implementation were not reviewed. (source)

Referenced by

Sources