arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10123cs.SEcs.CL

如果不是有缺陷的,就不要修复:关于使用LLM进行迭代式缺陷修复的动态研究

If It's Not Buggy, Don't Fix It: On the Dynamics of Iterative Bug-fixing with LLMs

Xietao Wang-Lin, Anton Isopoussu, Louis Mahon

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨LLM迭代式盲目修复缺陷的动态,发现其常误报无缺陷程序并陷入伪修复循环,机制探针揭示存在控制编辑倾向的引导向量,为自主修复系统提供见解。

中文摘要 AI 辅助

大型语言模型(LLM)在软件开发中已变得无处不在,基于LLM的自动化程序修复工具在代码审查中的应用也日益增多。在本报告中,我们探讨了将LLM作为缺陷修复器的迭代式盲目使用。在多种模型和修复环境中,我们发现LLM始终声称在完全无缺陷的程序中检测到缺陷,而修复有缺陷程序的成功率低于对正确程序造成的破坏率。我们还探讨了这一迭代过程的长期动态,发现它经常达到一种伪缺陷修复循环,即相同的更改被反复添加和移除,无限循环。最后,通过机制探针,我们揭示了存在一个控制编辑倾向的引导向量,这表明LLM具有对“有缺陷代码”的内部表征,而这种表征正是被错误激活以诱导伪缺陷修复的原因。这些结果为了解全自主缺陷修复系统的动态以及模糊目标下的停止条件提供了见解。

英文摘要

Large language models (LLMs) have become ubiquitous in software development, with LLM-based automated program repair tools increasingly used during code review. In this report, we explore the iterative blind use of LLMs as bug-fixers. Across multiple models and repair environments, we find that LLMs consistently claim to detect bugs in entirely bug-free programs while the rate of repair of buggy programs is less than that of the damage to correct programs. We also explore the long-term dynamics of this iterative process, and find that this frequently reaches a pseudo-bug-fixing cycle where the same changes are added and removed again ad infinitum. Lastly, via mechanistic probing, we unveil the existence of a steering vector which controls the editing propensity, suggesting that LLMs have an internal representation of ``buggy code", and that this representation is what is falsely activated to induce pseudo-bug fixing. These results provide insight towards the dynamics of fully autonomous bug-fixing systems, as well as stopping conditions under ambiguous goals.

发表机构

  • University of Warwick(华威大学)
  • UnlikelyAI

机构由 AI 辅助整理,请以论文原文为准。

↑