arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11593cs.SEcs.AI

编码智能体的可运行提交解缠

Runnable Commit Untangling for Coding Agents

Jinfeng Jiang, Dongsun Kim, Dayi Lin, Zhou Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对编码智能体生成的混乱提交补丁,提出保持代码可运行的RucTangle解缠方法与TangleEval评估框架,实验显示RucTangle可提升智能体修复退化的pass@1指标,证明软件工程实践在智能体时代的价值。

中文摘要 AI 辅助

编码智能体生成的大型混乱补丁会混合多个开发目的,导致代码难以审查和维护。提交解缠有望将此类大型补丁组织成解缠后、易于管理的提交。本文强调现有提交解缠研究的两个重要局限:其一,现有研究未考虑解缠后的提交具有顺序性,且应保持代码可运行;实践中,维护者不太可能接受导致代码无法运行的提交。其二,现有研究声称提交解缠有助于软件维护,但仅对解缠后的提交与开发者原始提交进行语法比较,未直接证明所声称的维护益处。为解决这些差距,本文做出两项新颖贡献:(1)RucTangle,首个在每次提交后保持代码可运行的智能体式提交解缠方法;(2)TangleEval,首个量化解缠后、易于管理的提交历史如何帮助编码智能体修复缺陷的评估框架。我们在131个智能体生成的补丁上,将RucTangle与四种解缠方法对比:RucTangle生成的所有历史均为可运行,而基线方法生成的提交历史中20.6%-37.4%不可运行。我们进一步收集453个引入退化(即导致此前通过的测试失败)的智能体生成补丁,并让另外两个编码智能体修复这些退化;用RucTangle生成的历史增强智能体上下文,使pass@1指标获得5.2%的绝对提升。我们还分析智能体轨迹,以了解它们如何利用解缠后的提交来导航和修复缺陷。我们的发现表明,在编码智能体时代采用成熟软件工程实践的价值,这拓宽了未来研究议程:智能体如何主动利用软件历史做出更好的开发决策?

英文摘要

Coding agents produce large, tangled patches that mix multiple development purposes, making the code hard to review and maintain. Commit untangling offers the promise of organizing such large patches into untangled, manageable commits. This paper emphasizes two important limitations in existing commit untangling studies. First, they do not consider that untangled commits are ordered and should leave the code runnable. In practice, maintainers are unlikely to accept commits that prevent the code from running. Second, existing studies claim that commit untangling helps software maintenance. However, they conduct syntactic comparisons between the untangled commits and developers' original commits without directly showing the claimed maintenance benefits. To address these gaps, this paper makes two novel contributions: (1) RucTangle, the first agentic method that untangles commits while keeping the code runnable after each commit; and (2) TangleEval, the first evaluation framework that quantifies how untangled, manageable commit histories help coding agents repair bugs. We compare RucTangle against four untangling methods on 131 agent-generated patches. All histories produced by RucTangle are runnable, while baselines produce 20.6%-37.4% unrunnable commit histories. We further collect 453 agent-generated patches that introduce regressions (i.e., causing previously passing tests to fail) and ask two other coding agents to repair regressions. Augmenting agent context with RucTangle-produced histories yields 5.2% absolute improvement in pass@1. We also analyze agent trajectories to learn how they use untangled commits to navigate and fix bugs. Our findings demonstrate the value of adopting established software engineering practices in the era of coding agents, which broaden the future research agenda: how can agents actively use software history to make better development decisions?

发表机构

  • Korea University(韩国大学)
  • University of Waterloo(滑铁卢大学)
  • University of Alberta(阿尔伯塔大学)
  • Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑