arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28887cs.SEcs.AIcs.LG

添加是机器,删除是人类:测量和缓解大语言模型代码编辑中的删除回避问题

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

Amir M. Ebrahimi, Mohammed Mehedi Hasan, Aaditya Bhatia, Gopi Krishnan Rajbahadur, Ahmed E. Hassan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大语言模型代码编辑中的删除回避问题,通过SWE-bench Verified等基准测量其表现,发现后期训练中加入删除相关训练可缓解该问题并提升代码编辑性能。

中文摘要 AI 辅助

大语言模型越来越多地用于编写和修复生产代码,但越来越多的证据表明,它们通过测试的补丁会让代码库更难维护。我们确定了一个具体原因:删除回避,即系统地倾向于保留预期编辑需要删除的代码的趋势。在官方SWE-bench Verified排行榜上的五个领先模型中,即使在所有五个模型都能解决的任务上,针对开发者补丁的删除召回率最高仅为71.7%;模型能找到超过92%所需删除的正确文件,但在不到52%的案例中删除了确切的行。相反,29.0%的通过补丁会将目标代码包装在防护或回退中,我们将这种模式称为“防护并执行”。这些补丁通过测试是因为原始测试很少检查删除:当我们对34个已验证任务进行改造,使其在目标代码保留时失败,四个涵盖闭源和开源权重的前沿模型的通过率从63.2%降至41.9%。由于实际修复会混合删除和添加,我们整理了CanItDelete,这是一个包含200个从真实提交中挖掘的任务的基准,其全部所需编辑都是删除。即使没有添加工作,最好的模型仍会在五分之一的任务中失败,而较小的开源模型则降至18.0%。然后,我们在四个累积提示下对GPT-5.6 Sol进行了消融研究;成功度几乎没有变化,直到我们提供确切的行,这几乎消除了不完整的删除,但仅将成功率提高到80.5%,因为模型随后会删除超出范围的内容或改为添加代码。最后,通过一项试点研究,我们展示了一个潜在的解决方案:在后期训练中教授删除可减少删除回避并提高更广泛的代码编辑性能,这表明该行为是训练不足而非无法实现的。

英文摘要

Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebases harder to maintain. We identify one concrete source: deletion avoidance, the systematic tendency to retain code that an intended edit requires removing. Across the five leading models on the official SWE-bench Verified leaderboard, deletion recall against the developer patch reaches at most 71.7% even on tasks all five solve, and models reach the right file for over 92% of required deletions but cut the exact line in under 52% of cases. Instead, 29.0% of passing patches wrap the targeted code in a guard or fallback, a pattern we call Guard-and-Go. Such patches pass because the original tests rarely check removal: when we retrofit 34 Verified tasks with tests that fail if the targeted code remains, four frontier models spanning closed and open weights fall from 63.2% to 41.9%. Because real repairs mix removal with addition, we curate CanItDelete, a benchmark of 200 tasks mined from real commits whose entire required edit is deletion. Even with the addition work gone, the best model still fails one task in five, and smaller open models fall to 18.0%. We then ablate GPT-5.6 Sol under four cumulative prompts; success moves little until we supply the exact lines, which nearly eliminate incomplete deletion yet raise success only to 80.5% because the model then deletes beyond the spans or adds code instead. Finally, through a pilot study we show one potential fix: teaching deletion during post-training reduces deletion avoidance and improves broader code-editing performance, suggesting the behavior is undertrained rather than beyond reach.

发表机构

  • Queen’s University(女王大学)

机构由 AI 辅助整理,请以论文原文为准。

↑