arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhoenixRepair:重新思考软件代理中的修复策略探索

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents

Tianyue Jiang, Yanlin Wang, Xin He, Daya Guo, Jiachi Chen, Ming Wen, Ensheng Shi, Xilin Liu, Yuchi Ma, Guanbin Li

arXiv 2607.18859首次发表:更新:

发表机构

Zhejiang University; Huazhong University of Science and Technology; Huawei CodeArts Model Team(浙江大学; 华中科技大学; 华为代码艺术模型团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有软件代理修复策略探索不足的问题,提出PhoenixRepair多代理框架,通过多位置采样、迭代反思优化等扩大搜索空间,实验表明该框架在解决率和故障定位精度上有提升,实现了7.8%的相对改进及76.0%的最高解决率Pass@1。

AI 中文摘要

虽然大语言模型极大地推动了自动问题解决,但现有基于代理的方法在修复策略探索方面存在根本局限性。这种不足体现在两个关键方面:一是对多个潜在编辑位置的探索有限,二是对每个位置修复尝试的探索也不足。为应对这些挑战,我们提出了PhoenixRepair,这是一个多代理框架,系统地探索多个候选编辑位置,并对补丁生成进行迭代反思和优化,从而扩大修复策略的搜索空间。我们的框架从多位置采样开始,可选择为困难任务添加基于图的定位信息,然后进行迭代反思和优化以生成更好的补丁,最终由所有历史尝试的提炼见解指导最终轮生成。在SWE-bench-Verified上的实验表明,PhoenixRepair在DeepSeek-V3.1下比SWE-agent实现了7.8%的最大相对改进,在MiniMax-M2.5下达到了76.0%的最高解决率Pass@1。同时,它比现有方法具有更高的故障定位精度。我们的代码可在此https URL获取。

英文摘要

While Large Language Models have greatly advanced automated issue resolution, existing agent-based methods exhibit a fundamental limitation in their insufficient exploration of repair strategies. This insufficiency manifests in two key aspects. First, the exploration of multiple potential edit locations is limited. Second, the exploration of repair attempts at each location is also insufficient. To address these challenges, we present PhoenixRepair, a multi-agent framework that systematically explores multiple candidate edit locations and performs iterative reflection and refinement on patch generation, thereby expanding the search space of repair strategies. Our framework begins with multi-location sampling, optionally augmented with graph-based localization information for difficult tasks, followed by iterative reflection and refinement to generate better patches, culminating in final-round generation guided by distilled insights from all historical attempts. Experiments on SWE-bench-Verified demonstrate that PhoenixRepair achieves the largest relative improvement of 7.8\% over SWE-agent under DeepSeek-V3.1, and attains the highest resolved rate of 76.0\% Pass@1 under MiniMax-M2.5. Meanwhile, it achieves higher fault localization accuracy than existing approaches. Our code is available at https://github.com/DeepSoftwareAnalytics/PhoenixRepair.

Comments14 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑