arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21997cs.SE

“回家吧,副驾驶,你喝醉了”:理解开发者对智能体生成的代码审查评论的回应

"Go Home Copilot, You're Drunk": Understanding Developer Responses to Agent-Generated Code Review Comments

Shamse Tasnim Cynthia, Ratnadira Widyasari, Banani Roy, Ting Zhang, David Lo

首次发表
浏览论文内容

中文总结 AI 辅助

研究开发者对智能体生成的代码审查评论的回应,分析五个智能体在342个Python仓库生成的54791条评论,探讨解决率、开发者经验及影响评论有用性的特征,发现不同智能体解决率有差异,还确定未解决评论原因及影响因素,为改进相关反馈提供见解。

中文摘要 AI 辅助

代码审查是软件工程开发中的关键质量保证实践,人工智能编码智能体越来越多地对拉取请求生成审查评论。然而,对于开发者如何实际回应此类智能体生成的反馈知之甚少。本文首次对智能体生成的代码审查评论的解决情况进行大规模实证研究。分析了GitHub上342个Python仓库中由五个广泛使用的编码智能体(Copilot、Cursor、Codex、Devin和Claude)生成的54791条评论。研究了智能体和评论类型的解决率、开发者经验的作用以及影响评论有用性的特征。结果表明,不同智能体的解决率差异很大,Copilot解决的评论占大多数(72.9%)。核心开发者解决了大部分智能体生成的反馈,特别是与“设计”和“可演化性”相关的评论,而外围开发者更多参与解决“功能缺陷”评论。通过对470条未解决评论讨论的开放式卡片分类,确定了十条解释评论未解决原因的讨论模式,“错误建议”和“故意设计决策”最为普遍。最后,分析表明内联“代码建议”的存在是评论解决的最强预测因素,而冗长复杂的评论不太可能被采纳。研究结果为改进人工智能生成的代码审查反馈及其融入开发工作流程提供了见解。

英文摘要

Code review is a critical quality assurance practice in software engineering development, and AI coding agents are increasingly generating review comments on pull requests. However, little is known about how developers actually respond to such agent-generated feedback. In this paper, we present the first large-scale empirical study on the resolution of agent-generated code review comments. We analyze $54{,}791$ comments generated by five widely used coding agents (i.e., Copilot, Cursor, Codex, Devin, and Claude) across $342$ Python repositories on GitHub. We examine (1) resolution rates across agents and comment types, (2) the role of developer experience, and (3) characteristics that influence comment usefulness. Our results show that resolution rate varies considerably across agents, with Copilot accounting for the majority of resolved comments (72.9\%). Core developers resolve the majority of agent-generated feedback, particularly for \textit{design} and \textit{evolvability}-related comments, while peripheral developers are more involved in resolving \textit{functional defect} comments. Through open card sorting of 470 unresolved comment discussions, we identify \textit{ten} discussion patterns explaining why comments remain unresolved, with \textit{incorrect suggestions} and \textit{intentional design decisions} being the most prevalent. Finally, our analysis reveals that the presence of an inline \textit{code suggestion} is the strongest predictor of comment resolution, while lengthy and complex comments are less likely to be acted upon. Our findings provide insights for improving AI-generated code review feedback and its integration into development workflows.

↑