arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExecCritic: 学会测试,测试以改进编码智能体

ExecCritic: Learn to Test, Test to Improve for Coding Agents

Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao

arXiv 2609.09133首次发表:更新:

发表机构

University of Wisconsin–Madison; Microsoft Research; Georgia Tech(威斯康星大学麦迪逊分校; 微软研究院; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ExecCritic通过分离测试生成与代码修复,结合角色特定强化学习,在SWE-bench Verified上将编码智能体解决率从61.2%提升至72.6%。

AI 中文摘要

执行反馈可以引导编码智能体对代码仓库进行正确的修复,但前提是测试能够捕捉到问题所请求的行为。智能体生成的测试可能编码不完整或不正确的行为目标;当同一条轨迹既编写补丁又编写测试时,它们的错误可能相互一致,从而产生虚假的信心。我们引入了ExecCritic,它将测试-验证-修订框架与一种针对角色特定的强化学习方案相结合,用于在该框架内训练智能体。该框架将测试构建与源代码修复分离:一个测试智能体独立生成仓库原生的测试,一个故障关闭的验证机制对测试进行资格验证并冻结它们,然后一个修复智能体根据这些测试的执行反馈来修订源代码,而不更改测试。两个角色均使用Qwen-3.5-35B-A3B作为骨干模型,并分别进行训练。在“学会测试”阶段,测试智能体学习生成行为上有效的测试,以区分正确和不正确的补丁。在“测试以改进”阶段,修复智能体同时学习直接的任务解决和反馈引导的修订。在SWE-bench Verified上,测试质量决定了反馈是否有帮助:保持基础修复智能体不变,来自基础测试智能体的测试将解决率从无测试基线的61.2%降低到57.3%,而来自GPT-5.6-sol的测试将其提高到65.3%。角色特定的后训练将Qwen测试智能体的Base-to-Gold成功率从22.2%提高到62.2%;将两个后训练的Qwen智能体组合起来达到72.6%,比原始无测试基线高出11.4个百分点,而在评估时无需更强模型或Oracle反馈。代码可在以下https URL公开获取。

英文摘要

Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair: a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from their execution feedback without changing the tests. Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In Learn to Test, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In Test to Improve, the Repair agent learns both direct task resolution and feedback-guided revision. On SWE-bench Verified, test quality determines whether feedback helps: holding the base Repair agent fixed, tests from the base Test agent reduce resolved rate from a no-test baseline of 61.2% to 57.3%, whereas tests from GPT-5.6-sol raise it to 65.3%. Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%; composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline without stronger-model or Oracle feedback at evaluation time. Code is publicly available at https://github.com/MSR-Orchard/execcritic.

Comments35 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑