arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18953cs.SE

去除噪声还是引入偏差?MSR过滤的隐藏代价

Removing Noise or Introducing Bias? The Hidden Cost of MSR Filtering

Mohit Kaushik, Jyoti Bawa

首次发表
浏览论文内容

中文总结 AI 辅助

本研究分析157万个GitHub仓库,揭示MSR研究中常用过滤标准引入的维护偏差、生态偏差和关系偏差,并建议采用分层抽样以减轻影响。

中文摘要 AI 辅助

像GitHub这样的开源软件平台是MSR研究的主要数据来源。由于该平台被从学生到开发者的不同用户广泛使用,并非所有仓库都是真正的工程项目。因此,为避免此类噪声,研究人员常应用多种筛选标准,而这些标准可能从根本上改变样本的人口统计学特征。为理解此类偏差,本研究旨在揭示源于这些任意阈值或标准的隐藏代价。我们分析了来自SEART平台的157万个仓库,并根据MSR研究中常用的阈值构建了多个数据集。我们识别出这些过滤过程中的维护偏差,该偏差掩盖了开源软件项目真实废弃率(73.42%)的现实。此外,这些策略偏向某些生态系统和治理风格。而且,采样策略还会扭曲变量之间的关系,导致关系偏差。因此,为避免此类偏差,研究人员应转向分层抽样并细化噪声检测标准。

英文摘要

OSS platforms like GitHub serve as a primary data source for MSR research. As the platform is widely used by different users, spanning from student to developer, not all repositories are actual engineered projects. Therefore, to avoid such noise, researchers often apply several criteria, which may fundamentally change the sample demographics. To understand such biases, this study aims to uncover the hidden cost originating from these arbitrary thresholds or criteria. We analyzed 1.57 million repositories from the SEART platform and constructed several datasets from the thresholds often applied in MSR research. We identify the maintenance bias in these filtering processes, which masks the true abandonment (73.42%) realities of OSS projects. Also, these strategies favor some ecosystems and governance styles. Moreover, the sampling strategy also distorts the relationship between variables, suffering from relational biases. Therefore, to avoid such biases, the researchers should shift towards stratified sampling and refining the criteria for noise detection.

发表机构

  • Guru Nanak Dev University(古鲁纳纳克德夫大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑