arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型如何读取错误报告?基于大语言模型的自动程序修复中注意力的实证研究

How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair

Ramtin Ehsani, Irene Manotas, Saurabh Pujar, Luca Buratti, Preetha Chatterjee

arXiv 2607.25873首次发表:更新:

发表机构

Drexel University; IBM Research(德雷塞尔大学; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于大语言模型的自动程序修复中模型注意力模式,通过分析319个真实错误,探究其在错误报告各部分的分布、成功与失败修复的注意力差异及与开发者重视信息的比较,发现注意力分配错误是失败关键因素,为改进系统提供见解。

AI 中文摘要

基于大语言模型(LLM)的自动程序修复系统发展迅速,但其性能仍不稳定。即使提供相同的上下文信息,LLM可能为一个错误生成正确补丁,却在另一个相关错误上失败。为何如此尚不清楚,模型如何优先处理错误报告中的各种信息以及模型注意力是否影响修复成功也不明确。本文首次对基于LLM的程序修复中的注意力模式进行实证研究,分析了来自SWE-bench Verified和Multi-SWE-bench的319个真实Python和Java错误,研究模型注意力在错误报告各部分的分布情况、成功与失败修复中各部分注意力模式的差异以及这些模式与开发者认为对错误修复重要的信息的比较。发现成功修复的特点是注意力分散在多个诊断组件上,而失败则常对版本信息等元数据过度局部关注。还观察到模型注意力与开发者确定的关键部分和短语的更强对齐与更高的修复成功率相关。结果提供了首个实证证据,即注意力分配错误是基于LLM的APR失败的关键因素,并为设计更具可解释性和可靠性的未来APR系统提供了可行见解。

英文摘要

Large Language Model (LLM)-based Automated Program Repair systems are advancing rapidly, yet their performance remains inconsistent. Even when provided with the same contextual information, an LLM may generate a correct patch for one bug but fail on another closely related bug. Why this happens remains poorly understood, and it is unclear how LLMs prioritize the diverse information in bug reports and whether model attention affects repair success. In this paper, we present the first empirical study of attention patterns in LLM-based program repair, providing interpretable insights into how models process bug reports and where their attention is concentrated during repair. We analyze 319 real-world Python and Java bugs from SWE-bench Verified and Multi-SWE-bench to study (RQ1) how model attention is distributed across bug report sections, (RQ2) how attention patterns within each section differ between successful and unsuccessful repairs, and (RQ3) how these patterns compare to information developers consider important for bug fixing. We find that successful repairs are characterized by diffused attention across multiple diagnostic components such as bug descriptions, stacktraces, and test cases, while failures often exhibit over-localized attention toward metadata such as version information. We further observe that stronger alignment between model attention and developer-identified key sections and phrases is associated with higher repair success. Our results provide the first empirical evidence that attention misallocation is a key factor in LLM-based APR failures, and offer actionable insights for designing more interpretable and reliable future APR systems.

CommentsAccepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE) 2026 Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑