发表机构
The University of Melbourne; Monash University(墨尔本大学; 莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
以CodeRabbit为例对智能代码审查进行实证研究,通过大量代码审查与开发者反馈数据,揭示其审查结果及被拒原因,还探索预测审查被拒的方法,指出改进空间与不足。
AI 中文摘要
智能代码审查(自主代理对拉取请求提供代码审查评论)日益融入开发工作流程,但开发者实际如何回应此类评论的实证证据有限。本文以CodeRabbit为例对智能代码审查进行实证研究。通过对239个GitHub仓库中10191个拉取请求的31073对代码审查和开发者反馈进行实证研究,结果表明智能审查的接受度不一:36.4%被接受,7.3%引发讨论,56.3%被拒绝。拒绝主要与误报、冗余或超出范围的无效建议以及与开发者意图和编码实践不一致有关。我们进一步发现,智能审查往往更关注功能问题而非与可演化性相关的评论,但它们更可能是无效的。为了提高审查实践的有效性,我们探索了各种基于大语言模型的方法来预测审查拒绝。我们发现,基于轻量级学习的方法可实现高达76%的F1分数,这表明代码审查与其相应反馈之间存在可学习的模式。我们的结果突出了CodeRabbit智能代码审查的当前状态,显示出改进的机会差距以及阻碍其有效性的缺点。
英文摘要
Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development workflows, yet there is limited empirical evidence on how developers respond to such comments in practice. In this paper, we present an empirical study of agentic code reviews using CodeRabbit as a case study. Through an empirical study of 31,073 pairs of code reviews and developer feedback from 10,191 pull requests across 239 GitHub repositories, our results show that agentic reviews receive mixed reception: 36.4% were accepted and 7.3% triggered discussion, while 56.3% were rejected. Rejections were primarily associated with invalid suggestions that were false positives, redundant, or out of scope, as well as misalignment with developer intent and coding practices. We further found that agentic reviews tend to focus more on functional concerns than evolvability-related comments, yet they were more likely to be invalid. To improve effectiveness in review practices, we explored various LLM-based approaches for predicting review rejection. We found that lightweight learning-based methods achieve up to 76% F1 score, suggesting learnable patterns exist between code reviews and their corresponding feedback. Our results highlight the current state of CodeRabbit's agentic code reviews, showing opportunity gaps for improvement, as well as shortcomings hindering its effectiveness.