发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型语言模型代码编辑中的过度编辑问题,构建评估框架,发现保留指令可缓解该问题,且强化学习能更好地学习最小编辑,明确编辑保真度是代码修复质量的独特维度。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于编辑现有代码,但仅正确性是不够的:有用的修复还应是最小化、可审查的,且忠实于原始实现。我们研究过度编辑,即模型重写代码超出修复错误所需的倾向。我们基于400个BigCodeBench问题构建评估框架,通过向参考解决方案注入受控的AST级损坏,为每个修复任务提供已知的最小补丁。在前沿LLMs中,过度编辑现象普遍存在,即使是像GPT-5.5这样的强大模型也不例外:高Pass@1指标可与不必要的大编辑和增加的认知复杂性共存。一条保留指令大幅减少了这种行为,将平均额外Levenshtein距离从0.195降至0.131,降低了26.6%的额外认知复杂性,并将Pass@1提高了2.3个百分点。然而,这些提升并非仅来自更大的推理预算或更大的模型。我们接下来探究是否能在训练后阶段直接学习最小编辑。我们观察到,监督微调会过拟合到已见的损坏模式,而强化学习则在域外编辑保真度与性能保持之间取得最佳权衡。这些结果将编辑保真度定位为代码修复质量的一个独特维度,并表明它可以被测量和学习。
英文摘要
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity. A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points. However, these gains do not simply follow from a larger reasoning budget or larger models. We next ask whether minimal editing can be learned directly during post-training. We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off. These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.
CommentsEMNLP 2026 (Main)