arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更好的理解,更好的修复?基于大语言模型的自动化程序修复中的幻觉研究

Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair

Xuemeng Cai, Jiakun Liu, Linhan Yang, Wei Ma, Lingxiao Jiang

arXiv 2609.04909首次发表:更新:

发表机构

Harbin Institute of Technology; Blekinge Institute of Technology; Singapore Management University(哈尔滨工业大学; 布莱金厄理工学院; 新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对基于LLM的自动化程序修复中的幻觉问题,通过多层分析、三项任务及832个Defects4J缺陷的实验,揭示修复与理解幻觉普遍存在,明确其表现、占比及成因。

AI 中文摘要

大型语言模型(LLM)已显著推动自动化程序修复(APR)的发展,但现有评估大多以结果为中心,对修复过程中的幻觉问题缺乏深入洞察。在APR中,幻觉不仅可能出现在最终补丁中,还可能出现在指导补丁生成的中间产物中。为填补这一空白,我们对APR全过程中的幻觉进行了多层分析,具体将幻觉定义为生成的补丁或中间产物未忠实基于可用修复证据。我们通过触发测试用例识别、行覆盖率预测、附加测试用例生成三项任务,分别研究最终补丁中的修复幻觉与中间产物中的理解幻觉,并在832个Defects4J缺陷上对三个代表性LLM进行自动评估与人工分析。结果显示,修复幻觉与理解幻觉均普遍存在:在所有模型与设置下,仅21.0%-55.9%的生成补丁通过开发者编写的测试套件;此外,尽管更准确的中间产物通常与成功修复相关,但这种关联并非总是成立。对812个抽样修复的人工分析发现,72.7%的案例存在修复幻觉,其中通过所有可用测试的补丁占比包含在内;不正确的因果定位和不正确的修复策略分别占这些幻觉的45.9%和18.5%。同时,模型频繁错误识别触发测试用例、错误预测涉及分支控制流的行覆盖率,以及生成缺少缺陷触发条件或预期行为错误的附加测试用例。

英文摘要

Large language models (LLMs) have significantly advanced automated program repair (APR), yet existing evaluations remain largely result-centric and provide limited insight into hallucination during repair. In APR, hallucination may arise not only in final patches but also in the intermediate artifacts that guide patch generation. To address this gap, we perform a multi-layered analysis of hallucination throughout the APR process. Specifically, we characterize hallucination as the production of patches or intermediate artifacts that are not faithfully grounded in the available repair evidence. We examine repair hallucination in final patches and understanding hallucination in intermediate artifacts through three tasks, namely triggering testcase identification, line coverage prediction, and additional testcase generation. We then evaluate three representative LLMs on 832 Defects4J bugs through automatic evaluation and manual analysis. Our results show that both repair and understanding hallucinations remain prevalent. Across models and settings, only 21.0%-55.9% of generated patches pass the developer-written test suite. Moreover, although more accurate intermediate artifacts are generally associated with successful repairs, this relationship does not always hold. Manual analysis of 812 sampled repairs identifies repair hallucinations in 72.7% of cases, including patches that pass all available tests; incorrect causal localization and incorrect repair strategies account for 45.9% and 18.5% of these hallucinations, respectively. Meanwhile, models frequently misidentify triggering testcases, mispredict line coverage involving branching control flow, and generate additional testcases with missing bug-triggering conditions or incorrect expected behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑