arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13292cs.SE

生成后优化:面向基于大语言模型的程序修复中正确且简洁的补丁

Refine After Generation: Toward Correct and Concise Patches in LLM-based Program Repair

Wenqiang Luo, Jacky Keung, Xiaoyu Shi, Yicheng Sun, Boyang Yang, Zhou Yang, Haoye Tian

首次发表
浏览论文内容

中文总结 AI 辅助

针对基于LLM的程序修复中补丁冗余问题,提出RECAP适配器,在保留或提升解决率的同时,大幅降低补丁大小,实现更优的大小-正确性权衡。

中文摘要 AI 辅助

大语言模型(LLM)已将自动程序修复(APR)推进到智能体系统可常规解决现实世界仓库级问题的阶段。然而,除了是否通过测试外,生成的补丁几乎未受到其他审查。本文中,我们将补丁冗余确定为基于LLM的APR中一个重大却被忽视的问题。在SWE-bench Verified上对28种最先进方法进行表征后,我们发现即使是成功的补丁也始终比开发者补丁更大、更复杂,其中值方法产生的总变更多121.78%,净变更多80.91%,圈复杂度高43.99%。我们进一步表明,这种冗余源于面向能力的设计选择,如迭代优化和宽泛上下文,且几乎无法通过输出格式或最小性提示等表面控制来减少。受这些发现的启发,我们提出生成后补丁优化,并设计了RECAP,这是一种轻量、即插即用的适配器,在生成后附加到现有修复框架。RECAP的优化器通过监督微调、直接偏好优化及从多源构建的补丁对数据集上的蒸馏推理轨迹进行训练。在四个宿主系统上,提示、提交解耦和最小性感知基线仅以牺牲49至217个已解决实例为代价减少补丁大小。相比之下,RECAP实现了显著更好的大小-正确性权衡,相对于开发者补丁,将平均总变更从+242.14%降至+4.24%,净变更从+348.24%降至-39.75%,同时保留或提升了最多42个实例的解决率。我们的结果表明,最小性不能简单地等同于句法压缩,将最小化与生成解耦为更易审查的修复提供了可行路径。

英文摘要

Large language models (LLMs) have advanced automatic program repair (APR) to the point where agentic systems routinely resolve real-world, repository-level issues. Yet the generated patch has received little scrutiny beyond whether it passes tests. In this paper, we identify patch verbosity as a major yet overlooked concern in LLM-based APR. Characterizing 28 state-of-the-art approaches on SWE-bench Verified, we find that even successful patches are consistently larger and more complex than developer patches, with the median approach producing 121.78% more total changes, 80.91% more net changes, and 43.99% higher cyclomatic complexity. We further show that this verbosity is rooted in capability-oriented design choices such as iterative refinement and broad context, and can hardly be reduced by surface-level controls such as output format or minimality prompts. Motivated by these findings, we formulate post-generation patch refinement and propose RECAP, a lightweight, plug-and-play adapter that attaches to existing repair frameworks after generation. RECAP's refiner is trained via supervised fine-tuning and direct preference optimization with distilled reasoning traces, on a dataset of patch pairs we construct from multiple sources. Across four host systems, prompting, commit-untangling, and minimality-aware baselines reduce patch size only by sacrificing 49 to 217 resolved instances. In contrast, RECAP achieves a substantially better size-correctness tradeoff, cutting average total changes from +242.14% to +4.24% and net changes from +348.24% to -39.75% relative to developer patches while preserving or improving resolution by up to 42 instances. Our results indicate that minimality cannot be simply reduced to syntactic compression, and that decoupling minimization from generation offers a practical path to more reviewable repairs.

↑