arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探索Linux内核中跨评审的语义稳定性

Exploring Semantic Stability Across Reviews in the Linux Kernel

Lucas Ciziks, Paulo Meirelles, Marco Aurélio Gerosa

arXiv 2608.10101首次发表:更新:

发表机构

Universidade de São Paulo; Northern Arizona University(圣保罗大学; 北亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以Linux内核IIO子系统为对象,提出函数级度量方法追踪10117条补丁函数轨迹,发现跨评审语义稳定性受无修改轨迹夸大,控制后仍有小残余效应,为相关研究提供初步探索。

AI 中文摘要

代码审查被认为会在补丁首次提交与最终落地版本之间大幅改变补丁的代码。然而,现有研究通常仅关注最终合并的补丁,未将其与首次提交的版本进行比较。我们提出一种函数级别的度量方法,追踪Linux IIO子系统补丁历史中的10117条轨迹(每条轨迹对应一个补丁系列编号修订版中的单个函数),并将相似度得分与不相关函数对的基线进行比较。直观来看,相似度近乎完全一致,但这在很大程度上是组合效应的结果:75.3%的被追踪轨迹在各版本间从未发生文本修改,贡献了100%的相似度,夸大了整体结果。限制到有实际修改的轨迹时,语义目的仍基本得以保留(平均相似度为0.990,而基线为0.909),但语义漂移主要集中在首次评审轮次,这主要是因为后续轮次包含更多无人触及的函数,而非编辑随时间变得更保守。控制该因素后,仍存在统计上可检测但较小的残余效应。这引出一个开放问题:近乎上限的相似度是反映目的的保留,还是测量工具无法检测到小型、局部编辑的重要性。我们将本研究作为初步探索,并概述后续步骤。

英文摘要

Code review is credited with substantially changing a patch's code between its first submission and the version that eventually lands. However, prior work typically studied only the final merged patch without comparing it to the first submission. We present a function-level measurement that tracks 10,117 trajectories (each function followed across the numbered revisions of one patch series) through the patch history of the Linux IIO subsystem, comparing similarity scores against unrelated function pairs as a baseline. A naive reading yields near-total similarity, but this is largely an artifact of composition: 75.3% of tracked trajectories are never textually modified between versions, contributing a trivial 100% similarity that inflates the headline. Restricting to the trajectories with a real edit, semantic purpose is still largely preserved (mean similarity 0.990 vs. a 0.909 baseline), but drift appears to concentrate in the first review round mainly because later rounds contain more functions that nobody touched, not because edits become more conservative over time. After controlling for it, a statistically detectable but small residual effect remains. This points to an open question: whether near-ceiling similarity reflects preserved purpose or a measurement tool that cannot detect the significance of small, localized edits. We present this work as a first look and outline next steps.

CommentsVEM 2026 - 14th Workshop on Software Visualization, Maintenance and Evolution

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑