AI 中文总结
本研究提出利用开源仓库拉取请求的提交构建复合提交解纠缠数据集,经验证其理想PR占比大幅提升,数据集规模远超以往,且可扩展至Python语言,为相关机器学习方法提供了更优训练数据。
AI 中文摘要
复合提交(Composite Commits, CC)是指将多个不相关变更打包为单个提交的形式,在软件开发中十分常见,会严重阻碍代码理解与维护。尽管已开发出基于机器学习的方法来将此类提交“解纠缠”为更小、连贯的变更集,但这些方法需要带有正确解纠缠标签的大规模训练数据,而准备此类数据集成本高昂,通常需要专家标注。本研究提出一种可扩展且经济高效的数据集构建方法,利用从开源仓库的拉取请求(Pull Requests, PRs)中提取的提交来构建数据集。我们对该数据集进行了实证验证,发现应用我们的过滤规则后,那些视为单个提交时存在纠缠、但功能分支上的每个单独提交均为原子性的理想PRs,占比从9.5%提升至55%。该复合提交数据集的规模是以往基于启发式方法构建的数据集的5.7倍以上。使用我们的新数据集,我们发现,即使考虑到我们提出的可能改变CC或STS大小的规则,基于PR的数据集与以往使用Herzig等人提出的启发式方法直接构建的数据集仍存在统计差异。在使用以往启发式方法构建数据集时,它们在影响投票者置信度的维度上存在统计差异,且可能对基于学习的方法产生影响。我们在原始Herzig等人的方法上验证了这种影响,该方法在我们的数据集上使用了投票者置信度。为证明我们的方法可扩展至其他语言,我们还创建了一个Python数据集并进行了实证验证,发现其理想PRs的占比(56.5%)与之前的结果相当。
英文摘要
Composite commits (CC), in which multiple unrelated changes are bundled into a single commit, are frequent in software development and significantly hinder code comprehension and maintenance. Although machine learning-based methods have been developed to ``untangle'' such commits into smaller, coherent change sets, these methods require large-scale training data with correct untangling labels. Preparing such datasets is costly and typically requires expert labelling. In this study, we propose a scalable and cost-effective method for dataset construction by leveraging commits extracted from open-source repositories' pull requests (PRs). We empirically validated our dataset and found that when applying our filtering rules, PRs that, when viewed as a single commit, are tangled, yet each individual commit on the feature branch is atomic (ideal PRs), increased from 9.5% to 55%. This composite commits dataset is more than 5.7 times larger than previous heuristic-based datasets. Using our new dataset, we find that the PR-based dataset differs statistically from previous datasets directly constructed using Herzig's proposed heuristics even after accounting for our proposed rules that may alter CC or STS sizes. When constructing datasets using the previous heuristics, they differ statistically along dimensions that impact the confidence voters and are likely to impact learning-based approaches. We validate the impact on the original Herzig \etal method, which used confidence voters across our dataset. To show that our approach extends to other languages, we also create a Python dataset which we empirically validate, finding comparable rates for ideal PRs (56.5%).