arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于编程的大型语言模型:实际修复还是重新实现错误代码?

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

Alexandru Stefan Stoica, Traian Rebedea, Marian Cristian Mihaescu

arXiv 2609.29410首次发表:更新:

发表机构

University Politehnica of Bucharest; University of Craiova(布加勒斯特理工大学; 克拉约瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究LLM修复代码缺陷时是否偏向重写而非局部修补,通过Codeforces数据集和多个GPT模型实验发现,LLM常过度修改代码,且从头生成比修补更有效,对AI辅助编程工具设计有重要启示。

AI 中文摘要

近期研究表明,大型语言模型(LLM)能够在包括竞争性编程在内的多种编程环境中有效解决问题并修复缺陷。现有方法主要独立评估LLM在问题解决或缺陷修复方面的性能,但未探讨这两种能力之间的关系。本研究聚焦于确定LLM在修复缺陷时偏离有缺陷解决方案的程度(相较于人类编写的补丁),以及其是否倾向于生成全新解决方案。我们构建了一个数据集,包含来自Codeforces数位用户的全部提交(约3000份),并将每个有缺陷的提交与其对应的人类修复版本进行匹配。以有缺陷解决方案与人类修复版本之间的相似度作为基线,我们在3个OpenAI GPT模型(gpt-5-nano、gpt-5-mini、gpt-5.1)上评估了LLM生成的缺陷修复质量。我们利用Codeforces-R1数据集(一个公开可用的数据集,包含由DeepSeek-R1模型生成的测试)来检查生成的解决方案是否解决了问题。我们的研究结果表明,与人类修复相比,LLM倾向于修改比必要更多的代码行,在某些情况下会生成全新的解决方案。我们还观察到,当允许LLM从头生成解决方案而非修补有缺陷的提交时,即使这些提交接近人类补丁,它们也能正确解决更多问题。这对AI辅助编程工具的设计具有重要意义,特别是在支持用户调试过程和促进增量式问题解决策略而非解决方案替换方面。

英文摘要

Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias towards generating entirely new solutions. We construct a dataset with all the submissions ($\sim$ 3000) from a couple of users from Codeforces, and we match each buggy submission with its corresponding human fix. By using the similarity between the buggy solution and the human fix as a baseline, we evaluate the quality of LLM-generated bug fixes on 3 OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1). We check if the generated solutions solve the problem by using the Codeforces-R1 dataset, an openly available dataset that has tests generated with the DeepSeek-R1 model. Our findings suggest that LLMs tend to modify more lines than necessary compared to human fixes and, in some cases, generate entirely new solutions. We also observe that LLMs solve more problems correctly when allowed to generate solutions from scratch rather than patch buggy submissions, even when those submissions are close to the human patch. This has important implications for the design of AI-assisted programming tools, particularly in supporting user debugging processes and promoting incremental problem-solving strategies rather than solution replacement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑