arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合并而非测量:编码代理修复的性能问题的实证研究

Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents

Zhenyu Qi, Haotang Li, Jinfu Chen, Huashan Chen, Yutong Zhao, Derui Zhu, Tomas Cerny, Bo Liu, Sen He

arXiv 2609.37985首次发表:更新:

发表机构

University of Arizona; Wuhan University; Institute of Information Engineering, Chinese Academy of Sciences; California State University, Long Beach; Rochester Institute of Technology(亚利桑那大学; 武汉大学; 中国科学院信息工程研究所; 加州州立大学长滩分校; 罗切斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析编码代理修复性能问题的合并情况,发现合并与否取决于代理在仓库的历史记录而非修复内容,且合并不保证修复达到声称的性能提升。

AI 中文摘要

编码代理开启的拉取请求(PR)声称能加速软件,但关于人类性能修复的研究很少说明维护者如何回应此类修复,或该声称是否成立。从AIDev v4的71,677个代理PR中,通过文本过滤和由语言模型及作者进行的编码手册编码,筛选出582个仓库中由六个代理修复的1,262个性能问题。我们对每个问题及其测试进行编码,并重新执行了23个被拒绝和30个被合并的修复。(1)57%的已关闭修复被合并,61%的拒绝未给出明确理由,且在主要基于代理构建的工作负载的三次运行试点中,23个重新执行的被拒绝声称中仅有6个成立。(2)接受率随代理在仓库中的过往记录(从31-37%升至70%)以及仓库在其其他代理PR上的预先合并率(从33%升至84%)而提高。合并的修复在其更改的行中删除的比例更大(0.26对0.15),这一差异在代理内和仓库内均成立,而在编码内容、描述、测试或测量中未检测到此类差异。(3)重复计算和冗余数据处理导致了44%的问题,46%的修复属于架构级别。(4)代理在37%的修复中更改了测试,11%携带性能测试或基准;在30个合并的修复中,18个达到了我们的交付标准,3个未达到声称,9个未显示显著增益或出现回退,14个在未测试输入上改变了行为。结果追踪的是仓库与代理的历史,而非修复的编码内容,且合并并不表明修复实现了其声称的效果。

英文摘要

Coding agents open pull requests (PRs) that claim to speed up software, but studies of human performance fixes say little about how maintainers respond to such a fix or whether its claim holds. From the 71,677 agent PRs of AIDev v4, a text filter and codebook coding by language models and by the authors select 1,262 performance issues fixed by six agents in 582 repositories. We code each issue and its tests and re-execute 23 rejected and 30 merged fixes. (1) 57% of closed fixes are merged, 61% of rejections give no stated reason, and only 6 of the 23 re-executed rejected claims held under our three-run pilot on mostly agent-built workloads. (2) Acceptance rises with the agent's track record in the repository (31-37% to 70%) and with the repository's pre-opening merge rate on its other agent PRs (33% to 84%). Merged fixes delete a larger share of the lines they change (0.26 versus 0.15), a difference that holds within agent and within repository, with no such difference detected in the coded content, description, tests or measurements. (3) Repeated computation and redundant data processing cause 44% of the issues, and 46% of fixes are architectural-level. (4) Agents change tests in 37% of fixes and 11% carry a performance test or benchmark; of the 30 merged fixes, 18 met our delivery criterion, 3 fell short of the claim, 9 showed no significant gain or regressed, and 14 change behavior on untested inputs. The outcome tracks the repository's history with the agent rather than the coded content of the fix, and a merge does not show that the fix delivers what it claims.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑