顺序重要吗?关于文件排序对代码评审有效性影响的实证研究
Does Order Matter? An Empirical Investigation into the Impact of File Ordering on Code Review Effectiveness
- University of Saskatchewan(萨斯喀彻温大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过大规模数据挖掘发现,文件排序影响代码评审中的注意力分配,字母顺序并非中立默认设置,并提出改进排序策略以提升评审有效性。
AI中文摘要:
现代代码评审是软件质量的核心,但其有效性取决于评审者的专业知识、变更特征以及工具展示变更的方式。大多数平台默认按字母顺序显示修改后的文件,尽管我们先前的研究表明,开发者发现这种排序在认知上与他们对多文件拉取请求的理解方式不一致。这些与排序相关的注意力模式是否在规模上影响结果仍不清楚。我们开展了一项关于文件排序与评审有效性的大规模研究,从182个使用五种编程语言的GitHub项目中挖掘了330,343个多文件拉取请求和756,814个文件实例。我们考察了文件位置是否与后续的缺陷修复变更相关,拉取请求大小是否调节这种关系,以及以评论为代表的评审者注意力是否与潜在缺陷结果一致。结果显示,文件位置、拉取请求大小、评审活动和潜在缺陷可能性之间存在统计显著但适度的关联。潜在缺陷率从位置1的56.7%上升到位置30的61.5%。拉取请求大小与潜在缺陷可能性呈非线性关系:约十个文件的拉取请求风险最低,而非常小和非常大的拉取请求风险较高。障碍模型显示,随着拉取请求大小的增长,注意力被稀释:每增加一个修改文件,收到任何评审评论的几率降低约8.7%。这些发现揭示了注意力-有效性差距:可见的评审活动并不一定能防止缺陷。因此,字母顺序排序并非中立的界面默认设置,而是一种结构性特征,塑造着注意力分配、评审覆盖范围以及对评审结果的信心。我们提出了上下文感知的文件排序、依赖感知的分组、风险感知的优先级排序以及逐文件覆盖指标,以使评审注意力更加可见和可操作。
英文摘要:
Modern code review is central to software quality, but its effectiveness depends on reviewer expertise, change characteristics, and how tools present changes. Most platforms display modified files alphabetically by default, although our prior work shows that developers find this ordering cognitively misaligned with how they understand multi-file pull requests. Whether these ordering-related attention patterns affect outcomes at scale remains unclear. We present a large-scale study of file ordering and review effectiveness, mining 330,343 multi-file pull requests and 756,814 file instances from 182 GitHub projects in five programming languages. We examine whether file position is associated with later bug-fixing changes, whether pull request size moderates this relationship, and whether reviewer attention, proxied by comments, aligns with latent bug outcomes. Results show statistically significant but modest associations among file position, pull request size, review activity, and latent bug likelihood. Latent bug rates rise from 56.7% at position 1 to 61.5% at position 30. Pull request size has a non-linear relationship with latent bug likelihood: pull requests of about ten files have the lowest risk, while very small and very large ones have elevated rates. Hurdle models show that attention is diluted as pull request size grows: each additional modified file reduces the odds of receiving any review comment by about 8.7%. These findings reveal an attention-effectiveness gap: visible review activity does not necessarily prevent defects. Alphabetical ordering is therefore not a neutral interface default, but a structural feature shaping attention allocation, review coverage, and confidence in review outcomes. We propose context-aware file ordering, dependency-aware grouping, risk-aware prioritization, and per-file coverage indicators to make review attention more visible and actionable.