arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29516cs.SEcs.AI

从代码审查到代码批判:大规模AI生成代码差异的意图、漂移与焦点

From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale

Chandra Maddila, Mashrur Rashik, Euna Mehnaz Khan, Smriti Jha, James Saindon, Nachi Nagappan, Peter C. Rigby

首次发表
浏览论文内容

中文总结 AI 辅助

针对AI生成代码超出传统审查能力且现有工具忽视核心问题的缺陷,提出ARCTIC系统,经实验验证其在意图预测、漂移检测等方面表现优异,可降低代码不一致性并获高认可。

中文摘要 AI 辅助

AI编码智能体生成代码的规模已超出传统同行审查的能力范围。与此同时,现有的AI代码审查工具过度关注风格、最佳实践等低价值建议,却忽视了人类审查员最关注的问题:正确性、安全性和性能。我们提出了ARCTIC,一个AI驱动的代码批判系统,它围绕三项能力重构代码审查:意图预测,即从对话日志和元数据中推断变更的原因;漂移检测,即通过反向翻译衡量开发者意图与智能体输出之间的差异;代码焦点,即对代码差异中最值得人类审查的区域进行排序。我们基于18000次代码审查得出的6主题分类法来支撑这些能力。离线评估显示,意图预测的F1值达到0.86,漂移检测与人类标注者的等级一致性接近完美(QWK=0.907),且在质量估计上,代码焦点的性能优于基线AI审查器2.4倍,同时使用的token数减少5倍。在实验推广中,漂移评分额外降低了5.76个百分点的代码不一致性(p=0.026),意图预测获得90.2%的认可,且自推出以来,自审查的代码差异未出现任何缺陷。

英文摘要

AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on the concerns human reviewers prioritize most: correctness, security, and performance. We present ARCTIC, an AI-powered Code Critique system that reframes code review around three capabilities: intent prediction, which infers why a change was made from conversation logs and metadata; drift detection, which measures divergence between the developer's intent and the agent's output via backtranslation; and code spotlight, which ranks the regions of a diff most warranting human scrutiny. We ground these capabilities in a six-theme taxonomy derived from 18,000 code reviews. Offline evaluation shows that intent prediction achieves 0.86 F1, drift detection reaches near-perfect ordinal agreement with human annotators (QWK = 0.907), and spotlight outperforms the baseline AI reviewer by 2.4x on quality estimation at 5x fewer tokens. In the experimental rollout, the drift scores reduces code misalignment by an additional 5.76 points (p = 0.026), intent prediction receives 90.2% approval, and zero defects have been attributed to self-reviewed diffs since launch.

发表机构

  • Concordia University(康考迪亚大学)
  • Meta

机构由 AI 辅助整理,请以论文原文为准。

↑