发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模比较智能体与人类编写的性能优化拉取请求,发现智能体PR合并率较低但策略相似,验证报告率相近,且证据常缺乏定量指标,凸显了建立严格多维评估标准的必要性。
AI 中文摘要
软件性能优化需要识别瓶颈和改进机会,并实证评估不同性能目标之间的效果与权衡。自主AI编码智能体向开源项目提交面向性能的拉取请求(PR),但这些变更是否反映了成熟的优化实践并充分证实其声称的价值,目前尚不清楚。我们扩展了针对407个性能PR的MSR 2026挖掘挑战试点研究,分析了截至2026年6月收集的1,130个智能体PR和1,130个人类编写的性能PR样本。我们比较了两组在采纳结果和补丁特征上的差异,并考察了它们的优化实践和验证行为。结果表明,智能体PR的合并率低于人类编写的PR(54.4%对73.3%),但两组使用了相似的优化策略,并以相近的比例报告验证。智能体PR更依赖静态推理,较少使用基准测试(46.1%对57.2%),尽管在研究期结束时基准测试的使用趋于一致。智能体PR对代码异味重构的验证报告频率与对资源定向变更的验证报告频率相同,而人类编写的PR则较少这样做。在两组中,约一半经过验证的PR未报告定量性能指标,且相关的权衡很少被量化。智能体优化实践已日益接近人类实践,但支持性证据仍不完整。确定智能体提出的优化是否得到严格、定量和多维证据的支持,仍然是一个重要挑战。这些发现促使评估基础设施和审查标准应衡量预期效果和相关成本,而不是仅将验证的存在视为充分条件。
英文摘要
Software performance optimization requires identifying bottlenecks and improvement opportunities, and empirically evaluating the effects and trade-offs across performance objectives. Autonomous AI coding agents submit performance-oriented pull requests (PRs) to open-source projects, yet it remains unclear whether these changes reflect established optimization practices and adequately substantiate their claimed value. Extending our MSR 2026 Mining Challenge pilot study of 407 performance PRs, we analyze a sample of 1,130 agentic and 1,130 human-authored performance PRs collected through June 2026. We compare adoption outcomes and patch characteristics between the two groups, and examine their optimization practices and validation behaviors. The results show that agentic PRs are merged less often than human-authored PRs (54.4% vs. 73.3%), but the two groups use similar optimization strategies and report validation at comparable rates. Agentic PRs rely more on static reasoning and less on benchmarks (46.1% vs. 57.2%), although benchmark use converges by the end of the study period. Agentic PRs also report validation for code-smell refactorings as often as for resource-targeting changes, whereas human-authored PRs do so less often. Across both groups, about half of validated PRs report no quantitative performance metric, and relevant trade-offs are rarely quantified. Agentic optimization practice has become increasingly similar to human practice, but the supporting evidence remains incomplete. Establishing whether agent-proposed optimizations are supported by rigorous, quantitative, and multidimensional evidence remains an important challenge. These findings motivate evaluation infrastructure and review criteria that measure intended effects and relevant costs rather than treating the presence of validation alone as sufficient.
Comments46 pages, 11 figures. Extends our MSR 2026 Mining Challenge paper. Huiyun Peng and Ricardo Calvo contributed equally. Replication package: https://github.com/PurdueDualityLab/EMSE-perf-pr-study