AI 中文总结
研究终端环境中智能代码审查的行为、成本等,基于审查者轨迹分析,发现智能审查者审查精度高但有探索等开销,成功审查与规划等有关,凸显轨迹感知和成本敏感评估对未来此类系统的好处。
AI 中文摘要
基于终端的环境中的智能代码审查能够在创建拉取请求之前的本地开发过程中提供早期反馈。然而,现有的评估仍然以性能为中心,未能捕捉基于仓库的智能审查者的动态行为。理解这些行为对于确定智能审查者在实践中如何成功、失败以及产生隐藏的运营成本至关重要。然后,我们基于审查者的轨迹分析他们的行为。我们的结果表明,智能审查者实现了更高的审查精度,但产生了大量的探索和验证开销,而成功的审查与更强的规划和更少的下游验证相关。这些发现凸显了对未来智能代码审查系统进行轨迹感知和成本敏感评估的潜在好处。
英文摘要
Agentic code review in terminal-based environments enables early feedback during local development before pull request creation. However, existing evaluations remain performance-centric and fail to capture the dynamic behaviors of repository-grounded agentic reviewers. Understanding these behaviors is critical for identifying how agentic reviewers succeed, fail, and incur hidden operational costs in practice. Then, we analyze the reviewers' behavior based on their trajectories. Our results show that agentic reviewers achieve higher review precision but incur substantial exploration and validation overhead, while successful reviews are associated with stronger planning and less downstream validation. These findings highlight the potential benefits of trajectory-aware and cost-sensitive evaluation of future agentic code review systems.