arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越词汇指标:AI代码审查中审查者习惯化的句子嵌入检测

Beyond Lexical Metrics: Sentence-Embedding Detection of Reviewer Habituation in AI Code Review

Haoran Yu, Lifei Liu, Danping Zhang

arXiv 2609.06213首次发表:更新:

发表机构

Nanchang Hangkong University(南昌航空大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过句子嵌入检测AI代码审查中审查者的习惯化现象,发现批准率上升但传统词汇指标无显著变化,而嵌入结构能有效捕捉适应信号,且语言变化滞后于行为变化。

AI 中文摘要

代码审查是AI生成代码与生产环境之间的关键质量检查点。随着AI编码代理大规模提交拉取请求,目前尚不清楚审查者是否会因重复接触而降低审查严格程度,以及审查评论是否能揭示这种变化。我们研究了来自400名重复审查者在207天内提交的11,429条审查记录,并配对了来自AIDev的10,104条人工撰写的内联评论。批准率从审查者早期阶段的30.5%上升到后期阶段的36.6%(Wilcoxon检验p = 8.6 x 10^-8;Cohen's d = 0.25)。然而,四个手工设计的语言特征——词汇多样性、香农熵、技术专业性和建设性可操作性——在暴露十分位数上均未显示单调下降(所有Spearman绝对rho <= 0.53,p >= 0.11;Bonferroni校正的Mann-Whitney检验p >= 0.36)。基于这些特征的逻辑回归分类器达到F1 = 0.485,低于多数类基线。句子嵌入结构确实携带信号:审查者后期评论的质心相对于早期质心的位移,大于在审查者内随机排列下的位移(Wilcoxon检验p < 0.001),并且使用三个嵌入统计量的小型MLP在5折审查者分层交叉验证下达到F1 = 0.74。格兰杰分析显示,批准率的变化可预测技术专业性的后续变化,在所有测试滞后下均显著(p < 0.001),而反向方向仅在四个滞后中的一个是显著的。因此,审查者的适应可在潜在分布结构中检测到,而非在经典词汇指标中,且语言变化跟随而非先于批准行为的变化。

英文摘要

Code review is a key quality checkpoint between AI-generated code and production. As AI coding agents submit pull requests at scale, it is unclear whether reviewers reduce scrutiny with repeated exposure and whether review comments reveal this change. We study 11,429 reviews from 400 repeat reviewers over 207 days, paired with 10,104 human-authored inline comments from AIDev. Approval rates rise from 30.5% in reviewers' early periods to 36.6% in late periods (Wilcoxon p = 8.6 x 10^-8; Cohen's d = 0.25). However, four hand-crafted linguistic features - lexical diversity, Shannon entropy, technical specificity, and constructive actionability - show no monotonic decline across exposure deciles (all Spearman absolute rho <= 0.53, p >= 0.11; Bonferroni-corrected Mann-Whitney p >= 0.36). A logistic-regression classifier based on these features reaches F1 = 0.485, below the majority-class baseline. Sentence-embedding structure does carry signal: reviewers' late-period comment centroids shift farther from their early-period centroids than under within-reviewer random permutations (Wilcoxon p < 0.001), and a small MLP using three embedding statistics reaches F1 = 0.74 under 5-fold reviewer-stratified cross-validation. Granger analysis shows that approval-rate changes predict later shifts in technical specificity at all tested lags (p < 0.001), while the reverse direction is significant at only one of four lags. Reviewer adaptation is therefore detectable in latent distributional structure rather than in classical lexical metrics, and language shifts follow rather than precede changes in approval behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑