arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14065cs.SEcs.AI

重新思考自动程序修复:缺陷复杂度、缺陷定位与大语言模型成本效率的影响

Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency

  • Colorado State University(科罗拉多州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, Fabio Santos

AI总结:

本研究通过实证分析,发现中等复杂度缺陷可超50%由低成本LLM修复,不精确缺陷定位会扩大APR技术差距,高成本LLM与强推理设置并非总能提升成本效率,DeepSeek-V3.2成本效率最优。

AI中文摘要:

背景:软件缺陷仍是开发中的关键挑战,需要有效的自动程序修复(APR)技术。基于大语言模型(LLM)的APR系统虽已展现潜力,但现有研究主要关注整体修复效果,缺陷复杂度、缺陷定位、推理设置及修复成本效率的影响仍未得到充分探索。目标:本研究针对基于LLM的APR开展全面实证分析,重点探究缺陷复杂度、缺陷定位、推理设置及成本如何影响修复性能。方法:我们使用三种LLM(DeepSeek、GPT和Llama)评估两种APR技术(ChatRepair和CodeCorrector),并通过多维度实证框架与统计分析,考察其在不同缺陷复杂度水平及定位策略下的性能。结果:尽管结构复杂的缺陷与不精确的缺陷定位会增加修复难度,但基于LLM的APR技术仍能达到具有竞争力的修复效果;不精确的缺陷定位会显著扩大APR技术间的性能差距;此外,更高成本的LLM与更强的推理设置并不总能带来更优的成本效率,揭示了修复效果与计算成本之间存在不容忽视的权衡关系。结论:超过50%的中等复杂度缺陷可由低成本基于LLM的APR技术修复;随着缺陷定位精度降低,APR技术间的修复效果差距会扩大;GPT-5比DeepSeek-V4-pro和DeepSeek-V3.2分别多修复7个和39个复杂缺陷,而DeepSeek-V3.2的总修复成本展现出最优的成本效率表现。

英文摘要:

Background: Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. While Large Language Model (LLM)-based APR systems have shown promise, prior studies primarily focus on overall repair effectiveness. The effects of bug complexity, fault localization, reasoning settings, and repair cost-effectiveness remain insufficiently explored. Aims: This study presents a comprehensive empirical analysis of LLM-based APR, focusing on how repair performance is shaped by bug complexity, fault localization, reasoning settings, and costs. Method: We evaluate two APR techniques (ChatRepair and CodeCorrector) using three LLMs (DeepSeek, GPT, and Llama), and examine their performance across diverse levels of bug complexity and localization strategies through a multi-dimensional empirical framework and statistical analysis. Results: Although structurally complex bugs and imprecise fault localization make repair more challenging, LLM-based APR techniques still achieve competitive repair effectiveness. Imprecise fault localization can substantially enlarge the performance gap between APR techniques. Furthermore, higher-cost LLMs and stronger reasoning settings do not consistently yield better cost-efficiency, revealing a nontrivial trade-off between repair effectiveness and computational cost. Conclusions: Over 50% of moderately complex bugs can be repaired by low-cost LLM-based APR techniques. The repair effectiveness gap between APR techniques becomes larger as fault localization becomes less precise. GPT-5 repairs 7 and 39 more complex bugs than DeepSeek-V4-pro and DeepSeek-V3.2, respectively; whereas the total repair cost of DeepSeek-V3.2 shows the best cost-efficiency performance.

补充信息

↑