arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Opera:面向长时程编码智能体的口头批评框架

Opera: A Verbal Critic Framework for Long-horizon Coding Agents

Kai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan, Yifan Zhang, Dingjie Song, Dimitris N. Metaxas, Silvio Savarese, Ran Xu, Zeyuan Chen

arXiv 2609.33987首次发表:更新:

AI 中文总结

Opera 提出一种口头批评框架,通过持久化纠正、事件驱动审查和反馈审计,提升长时程编码智能体的解决率,并利用其轨迹数据微调模型,在多个基准上取得显著改进。

AI 中文摘要

长时程编码智能体需要及时的纠正,然而当反馈误判正在进行的工作或未能解决根本问题时,反馈可能无效甚至有害。现有的批评者侧重于评估轨迹并生成反馈,但很少跟踪反馈交付后发生的情况。我们提出了 Opera,一个口头批评框架,它将每次纠正视为一条持久记录,并持续跟进直至诊断出的问题得到解决。Opera 通过周期性和事件驱动的触发器决定何时进行审查,使用类型化算子诊断问题,在交付前根据可见证据审计反馈,并跟踪智能体后续行动以区分单纯服从与真正解决。作为测试时的批评者,Opera 在 Terminal-Bench 2.1、SWE-Bench Pro 子集和 DeepSWE v1.1 上,跨四个策略模型,将非批评智能体的解决率分别提高了最多 12.4、15.0 和 8.9 个百分点,并在所有三个基准上取得了竞争性批评基线中最高的平均解决率,同时当策略自我批评时也能改进策略模型。超越推理阶段,Opera 引导的轨迹提供了近似策略的训练数据:在此数据上微调 Qwen3.5-9B,在保留的 SWE-Bench Pro 仓库上,其解决率在推理时无批评者的情况下提高了 10.2 个百分点,与在更强模型轨迹上微调的效果相当,同时在切换框架(即从 Openhands 切换到 Terminus-2)时保持性能,而后者则大幅下降。

英文摘要

Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades. Our code is available at: https://github.com/dongyuanjushi/Opera.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑