arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13196cs.SE

从以人类为中心到智能体代码审查:不同代际生成式人工智能技术对审查质量的影响

From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality

Suzhen Zhong, Shayan Noei, Bram Adams, Ying Zou

首次发表
浏览论文内容

中文总结 AI 辅助

研究不同代际生成式人工智能技术对代码审查质量的影响,通过分析102万条GitHub拉取请求,识别三种人工智能审查者采用模式,将审查讨论建模为交互序列,发现智能体协作模式虽能提高效率,但未提升质量,为设计高效代码审查过程提供实证指导。

中文摘要 AI 辅助

代码审查有助于在代码集成前维护软件质量,但给人工审查者带来巨大工作量。随着生成式人工智能融入软件开发,代码审查正从主要由人工审查转向人工智能支持的审查过程,包括大语言模型(LLM)审查者和人工智能智能体审查者与人工审查者共同参与。然而,我们仍缺乏关于这种转变如何影响审查效率和质量的实证证据。本文研究了来自207个GitHub项目的102万条审查过的拉取请求,这些项目跨越了三个代码审查时代:以人类为中心的审查、LLM辅助审查和智能体代码审查。我们识别出三种人工智能审查者采用模式:逐步采用人工智能、快速采用LLM和快速采用人工智能智能体。我们进一步将拉取请求审查讨论建模为审查者交互序列,以描述人工、LLM和人工智能智能体审查者在审查过程中的协作方式。结果表明,在逐步采用人工智能和快速采用人工智能智能体模式下,涉及智能体的协作模式,特别是由人工智能智能体发起或涉及多个人工智能智能体的审查,与更快的审查决策相关。然而,这些效率提升并未转化为更好的审查质量。我们还发现,审查活动和拉取请求类型在各个时代都很重要,而一旦LLM和人工智能智能体审查者参与,人工与人工智能的协作模式就成为审查效率的最强解释因素。这些发现为设计在不削弱审查质量的情况下提高效率的人工智能支持的代码审查过程提供了实证指导。

英文摘要

Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software development, code review is shifting from a primarily human review process toward AI-supported review processes in which large language model (LLM) reviewers and AI agent reviewers participate alongside human reviewers. However, we still lack empirical evidence on how this transition affects review efficiency and review quality. In this paper, we study 1.02 million reviewed pull requests from 207 GitHub projects that transition across three code review eras: human-centric review, LLM-assisted review, and agentic code review. We identify three AI reviewer adoption practices: Gradual AI Adoption, Rapid LLM Adoption, and Rapid AI Agent Adoption. We further model pull request review discussions as reviewer interaction sequences to characterize how human, LLM, and AI agent reviewers collaborate during the review process. Our results show that agent-involved collaboration patterns, especially reviews initiated by AI agents or involving multiple AI agents, are associated with faster review decisions under Gradual AI Adoption and Rapid AI Agent Adoption. However, these efficiency gains do not translate into better review quality. We also find that review activity and pull request type remain important across eras, while human-AI collaboration patterns become the strongest explanatory factor for review efficiency once LLM and AI agent reviewers participate. These findings provide empirical guidance for designing AI-supported code review processes that improve efficiency without weakening review quality.

↑