arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用自我增强的ChatGPT进行自动程序修复的泛化性研究:一项实证研究

Towards the Generalizability of Leveraging ChatGPT in APR via Self-enhancing: An Empirical Study

Qingyuan Li, Chuanyi Li, Yaopeng Yang, Ziwen Ge, Jidong Ge, Bin Luo

arXiv 2609.22844首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology, Nanjing University(南京大学软件新技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过三个基准和两个模型实证评估ChatGPT增强APR方法的泛化性,发现收益因基准和模型而异,直接提供修复特定外部信息可能更有效。

AI 中文摘要

自动程序修复(APR)越来越依赖于大型语言模型(LLMs)。ChatGPT增强的APR使用自我纠正和自主智能体等技术,在不修改模型参数的情况下改进修复效果。尽管这些方法在Defects4J和SWE-bench上报告了强劲的结果,但增强收益在不同基准上的稳定性仍未得到充分探索。我们评估了三种ChatGPT增强的APR方法在三个具有代表性且长期存在的基准上的表现。使用GPT-3.5-Turbo时,SRepair在HumanEval-Java上比在Defects4J上获得了更大的绝对收益,而SRepair和FixAgent在没有外部信息的情况下在BugsInPy上产生了负收益。使用GPT-5.4-mini时,所评估的方法在Defects4J上比在HumanEval-Java上获得了更大的绝对收益,而在BugsInPy上的收益为非负但有限。我们通过代码转换和基准特定微调来研究基准相关因素。代码转换减少了Defects4J上的增强收益,而基准特定微调增加了BugsInPy上的收益。在BugsInPy上,直接向GPT-3.5-Turbo提供错误消息和触发测试所产生的正确修复数量多于所评估的ChatGPT增强APR方法。这些发现强调了使用多种指标评估跨基准和模型的泛化性的必要性,并表明当增强方法的收益有限时,直接提供修复特定的外部信息可能比增强方法更有效。

英文摘要

Automated Program Repair (APR) increasingly relies on Large Language Models (LLMs). ChatGPT-enhanced APR uses techniques such as self-correction and autonomous agents to improve repair without modifying model parameters. Although these approaches report strong results on Defects4J and SWE-bench, the stability of enhancement gains across benchmarks remains under-explored. We evaluate three ChatGPT-enhanced APR methods on three representative, long-standing benchmarks. With GPT-3.5-Turbo, SRepair achieves a larger absolute gain on HumanEval-Java than on Defects4J, while SRepair and FixAgent without extrinsic information yield negative gains on BugsInPy. With GPT-5.4-mini, the evaluated methods achieve larger absolute gains on Defects4J than on HumanEval-Java, while gains on BugsInPy are non-negative but limited. We investigate benchmark-related factors through code transformations and benchmark-specific fine-tuning. Code transformations reduce enhancement gains on Defects4J, while benchmark-specific fine-tuning increases gains on BugsInPy. Directly supplying GPT-3.5-Turbo with error messages and triggering tests yields more correct repairs than the evaluated ChatGPT-enhanced APR methods on BugsInPy. These findings highlight the need to evaluate generalizability across benchmarks and models using multiple metrics, and suggest that directly providing repair-specific extrinsic information may be more effective than enhancement methods when their gains are limited.

Comments18 pages, 8 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑