$ cat wiki/papers/2026/2609.33987-opera.md
Opera — persistent feedback for coding agents
TL;DR
Opera keeps a critic’s correction active until the underlying problem is resolved. Its abstract reports gains during inference and after training on critic-guided rollouts. (source)
Authors & Org
Kai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan, Yifan Zhang, Dingjie Song, Dimitris N. Metaxas, Silvio Savarese, Ran Xu and Zeyuan Chen. Affiliations are not stated in the captured abstract page. First submitted 2026-09-27; the reviewed version was revised 2026-10-08. (source)
Method
Periodic and event-driven triggers initiate review. Typed diagnostic operators identify problems, an evidence audit checks feedback before delivery, and persistent notes track subsequent actions to distinguish compliance with advice from resolution of the problem. The authors also use these approximately on-policy rollouts for fine-tuning. (source)
Results
- Across four policy models, the authors report resolve-rate gains of up to 12.4 percentage points on Terminal-Bench 2.1, 15.0 on a SWE-Bench Pro subset, and 8.9 on DeepSWE v1.1. The abstract also reports gains when a model critiques itself. (source)
- Fine-tuning Qwen3.5-9B improves held-out SWE-Bench Pro resolve rate by 10.2 percentage points without an inference-time critic. The authors report matching stronger-model rollout training while retaining performance when switching from Openhands to Terminus-2; the comparator degrades under that switch. (source)
Significance
For Agents (LLM Agents), this offers an explicit record of unresolved corrections, complementing the outstanding-requirement record in RunningTab — tracking unfinished workspace requirements. Interpretation: generating useful advice and checking whether it fixed the problem are separate responsibilities. (source) (RunningTab)
Open Questions
The abstract does not provide absolute resolve rates, per-model rows, or a cost-matched comparison. The reported maxima should not be read as gains for every model. Full paper and implementation were not reviewed. (source)
Cite
arXiv:2609.33987 · Code · (source)