arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17598cs.SE

并非所有智能体都生而平等:五个自主编码智能体在真实场景中的代码质量与合并后维护

Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild

  • Texas Tech University(德克萨斯理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Obada Kraishan

AI总结:

本研究分析五个商业编码智能体的3.7万余个真实PR,发现代码质量因供应商而异,Codex回滚率低于人类,Devin高于人类,且智能体代码安全异味更少,审查负担分布不均。

AI中文摘要:

自主编码智能体现在以两年前无法达到的规模在公共仓库中开启拉取请求,但关于这些代码在合并后发生的情况却鲜为人知。本文研究了来自五个商业智能体(OpenAI Codex、Devin、GitHub Copilot、Cursor 和 Claude Code)的 37,623 个带来源标记的拉取请求(PR),以及一个匹配的人类基线,这些数据取自 2024 年 12 月至 2025 年 7 月间的 2,807 个 GitHub 仓库。我们将 AIDev 数据集与 58,792 个缓存的 GitHub API 响应相结合,以衡量新增代码中的安全异味、结构可维护性、合并后变动、回滚率以及人类审查行为。三个结果尤为突出。首先,质量差异因供应商而异,而非统一:Codex 撰写的 PR 被回滚的频率约为人类 PR 的一半(6.1% 对比 11.5%,比值比为 0.50),而 Devin 的 PR 被回滚的频率更高(14.5%,比值比为 1.31)。其次,跨供应商汇总的智能体代码比人类代码更不可能包含安全异味(比值比为 0.63),这主要归因于硬编码凭据和 eval 风格结构的减少。第三,审查工作分布不均:Copilot 的 PR 吸引了最多的人类审查和变更请求,而 Claude Code 的 PR 等待首次人类审查的时间最长(中位数为 12.6 小时)。所有流水线代码、统计报告和图表均已发布以供复现。

英文摘要:

Autonomous coding agents now open pull requests in public repositories at a scale that was out of reach two years ago, yet little is known about what happens to that code after it lands. This paper studies 37,623 provenance-labeled pull requests (PRs) from five commercial agents (OpenAI Codex, Devin, GitHub Copilot, Cursor, and Claude Code) and a matched human baseline, drawn from 2,807 GitHub repositories between December 2024 and July 2025. We combine the AIDev dataset with 58,792 cached GitHub API responses to measure security smells in added code, structural maintainability, post-merge churn, revert rates, and human review behavior. Three results stand out. First, quality differences are vendor-specific rather than uniform: Codex-authored PRs were reverted about half as often as human PRs (6.1% vs. 11.5%, odds ratio 0.50), while Devin PRs were reverted more often (14.5%, odds ratio 1.31). Second, agent code pooled across vendors was less likely than human code to contain a security smell (odds ratio 0.63), driven by fewer hardcoded credentials and eval-style constructs. Third, review effort concentrates unevenly: Copilot PRs drew the most human reviews and change requests, and Claude Code PRs waited the longest for a first human review (median 12.6 hours). All pipeline code, statistical reports, and figures are released for replication.

补充信息

↑