发表机构
State Key Laboratory for Novel Software Technology, Nanjing University(南京大学软件新技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过三个基准和两个模型实证评估ChatGPT增强APR方法的泛化性,发现收益因基准和模型而异,直接提供修复特定外部信息可能更有效。
AI 中文摘要
自动程序修复(APR)越来越依赖于大型语言模型(LLMs)。ChatGPT增强的APR使用自我纠正和自主智能体等技术,在不修改模型参数的情况下改进修复效果。尽管这些方法在Defects4J和SWE-bench上报告了强劲的结果,但增强收益在不同基准上的稳定性仍未得到充分探索。我们评估了三种ChatGPT增强的APR方法在三个具有代表性且长期存在的基准上的表现。使用GPT-3.5-Turbo时,SRepair在HumanEval-Java上比在Defects4J上获得了更大的绝对收益,而SRepair和FixAgent在没有外部信息的情况下在BugsInPy上产生了负收益。使用GPT-5.4-mini时,所评估的方法在Defects4J上比在HumanEval-Java上获得了更大的绝对收益,而在BugsInPy上的收益为非负但有限。我们通过代码转换和基准特定微调来研究基准相关因素。代码转换减少了Defects4J上的增强收益,而基准特定微调增加了BugsInPy上的收益。在BugsInPy上,直接向GPT-3.5-Turbo提供错误消息和触发测试所产生的正确修复数量多于所评估的ChatGPT增强APR方法。这些发现强调了使用多种指标评估跨基准和模型的泛化性的必要性,并表明当增强方法的收益有限时,直接提供修复特定的外部信息可能比增强方法更有效。
英文摘要
Automated Program Repair (APR) increasingly relies on Large Language Models (LLMs). ChatGPT-enhanced APR uses techniques such as self-correction and autonomous agents to improve repair without modifying model parameters. Although these approaches report strong results on Defects4J and SWE-bench, the stability of enhancement gains across benchmarks remains under-explored. We evaluate three ChatGPT-enhanced APR methods on three representative, long-standing benchmarks. With GPT-3.5-Turbo, SRepair achieves a larger absolute gain on HumanEval-Java than on Defects4J, while SRepair and FixAgent without extrinsic information yield negative gains on BugsInPy. With GPT-5.4-mini, the evaluated methods achieve larger absolute gains on Defects4J than on HumanEval-Java, while gains on BugsInPy are non-negative but limited. We investigate benchmark-related factors through code transformations and benchmark-specific fine-tuning. Code transformations reduce enhancement gains on Defects4J, while benchmark-specific fine-tuning increases gains on BugsInPy. Directly supplying GPT-3.5-Turbo with error messages and triggering tests yields more correct repairs than the evaluated ChatGPT-enhanced APR methods on BugsInPy. These findings highlight the need to evaluate generalizability across benchmarks and models using multiple metrics, and suggest that directly providing repair-specific extrinsic information may be more effective than enhancement methods when their gains are limited.
Comments18 pages, 8 figures, 9 tables