AI 中文总结
该研究构建大规模 AI 对 AI 代码审查数据集,分析 GitHub 上 AI 智能体生成 PR 的审查情况,发现跨产品审查占比约 1.6%且增速快,审查输出因配置而异,闭环 AI 对 AI 审查呈增长趋势但占比仍低。
AI 中文摘要
AI 编码智能体正日益融入软件开发工作流,在拉取请求(PR)流程的两侧运行:AI 创作智能体创建或修改 PR,AI 审查智能体对其进行评估。这形成了一个闭环,即一个 AI 编码智能体审查另一个智能体生成的贡献。我们通过将 AI 生成的 PR 与来自 CodAGE(一个公开的编码智能体生成的 GitHub 事件数据集)中的 AI 生成的审查事件相关联,构建了一个大规模的 AI 对 AI 代码审查数据集。该数据集包含 248641 个独特的 AI 生成 PR,每个 PR 至少收到一次 AI 生成的审查;其中 45269 个 PR 收到跨产品审查,208145 个收到同产品审查,4773 个 PR 同时收到两种审查。跨产品 AI 对 AI 审查约占已识别智能体生成 PR 的 1.6%,但绝对规模可观,且从 2025 年第一季度到 2025 年第三季度,其数量增长了两个数量级以上。审查输出因创作者-审查者配置而异:CodeRabbit 对 Claude Code 生成的 PR 的评论中,35.0% 被标记为重构评论,而对 Copilot 生成的 PR 的这一比例为 10.5%,不过这一差异可能反映了 PR 的特征而非审查者的特征。在四个双角色审查者中,有三个在同产品组中每个 PR 的平均评论数高出 58%-65%,尽管效应量较小或可忽略,且差异集中在上尾部分。在具有完整非负时间戳的配对中,跨产品配对的观察到的中位延迟为 1.2 分钟,而同产品配对为 4.7 分钟;时间戳可用性的差异和审查者构成限制了该比较。总体而言,闭环 AI 对 AI 审查正在增加,但在已识别的智能体活动中仍占少数,审查输出因创作智能体组和产品配置而异。
英文摘要
AI coding agents are increasingly integrated into software development workflows, operating on both sides of the pull-request (PR) process: AI authoring agents create or modify PRs, while AI reviewers evaluate them. This creates a closed loop in which one AI coding agent reviews a contribution attributed to another. We construct a large-scale dataset of AI-to-AI code review by linking AI-attributed PRs with AI-attributed review events from CodAGE, a public dataset of coding-agent-generated GitHub events. Our dataset contains 248,641 unique AI-attributed PRs that received at least one AI-attributed review. Of these, 45,269 received cross-product review and 208,145 received same-product review; 4,773 PRs received both. Cross-product AI-to-AI review occurred in approximately 1.6% of identified agent-authored PRs but was substantial in absolute terms, and its volume increased by more than two orders of magnitude from 2025-Q1 to 2025-Q3. Reviewer output varied across author-reviewer configurations. CodeRabbit labeled 35.0% of its comments on Claude Code-authored PRs as refactor comments, compared with 10.5% on Copilot-authored PRs, although this difference may reflect characteristics of the PRs rather than the reviewer. For three of four dual-role reviewers, mean comments per PR were 58-65% higher in the same-product group, although effect sizes were small or negligible and the difference was concentrated in the upper tail. Among pairs with complete, nonnegative timestamps, the observed median latency was 1.2 minutes for cross-product pairs and 4.7 minutes for same-product pairs; differential timestamp availability and reviewer composition limit this comparison. Overall, closed-loop AI-to-AI review is increasing but remains a minority of identified agent activity, with review output varying across authoring-agent groups and product configurations.
CommentsAccepted at the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026), Emerging Results, Vision, and Reflection Papers Track