arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开源项目中不完整软件变更经验数据集的构建

The Construction of an Empirical Dataset of Incomplete Software Changes from Open Source Projects

Savira Ramadhanty, Profir-Petru Pârţachi, Yoshiya Ishida, Takashi Kobayashi

arXiv 2609.34317首次发表:更新:

发表机构

Institute of Science Tokyo(东京科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建了从开源项目中挖掘真实不完整变更的数据集,分析其特征,并重新评估了共同变更规则提取方法LCExtractor,发现多数遗漏变更涉及少量文件。

AI 中文摘要

在软件开发过程中,对软件组件的修改可能会在系统中传播,需要精确识别并正确修订所有受影响的组件。这是一项复杂的任务,开发者常常(45.7%)遗漏相关变更。为解决此问题,已有多种方法从修订历史中频繁共同变更的文件中提取共同变更规则。然而,以往研究使用人为制造的不完整变更来评估这些方法,这可能无法代表真实世界的数据。为解决此问题,我们通过从一系列开源软件中挖掘不完整变更来构建数据集,利用问题跟踪平台中关于引入缺陷及其相应修复的信息。我们还利用该数据集分析不完整变更的特征,发现89.4%的遗漏变更涉及五个或更少的文件。最后,我们在所构建的数据集上重新评估了现有的共同变更规则提取方法LCExtractor,并确定了最优排序标准及所用提交数量的影响。

英文摘要

During software development, a modification to a software component may propagate across the system, requiring precise identification and correct revision of all affected components. This is a complex task, and developers often (45.7%) miss related changes. To address this, several methods have been developed to extract co-change rules from files that are frequently changed together in the revision history. However, previous research evaluated the methods using artificially created incomplete changes, which may not be representative of real-world data. To solve this problem, we construct a dataset by mining incomplete changes from a collection of open-source software, using information about induced bugs and their respective fixes from an issue tracking platform. We also analyze the characteristics of incomplete changes using this constructed dataset and found that 89.4% of missed changes involved five or fewer files. Finally, we re-evaluate LCExtractor, an existing co-change rule extraction method, on our constructed dataset, and we identify the optimal sorting criterion and the impact of the number of used commits.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑