机器学习研究软件中用于可复现性保障的突变测试:一项实证研究
Mutation Testing for Reproducibility Safeguards in Machine Learning Research Software: An Empirical Study
浏览论文内容
中文总结 AI 辅助
本研究以MLReproMutate为工具,对39个机器学习研究仓库开展实证研究,发现现有验证工作流仅能检测8.7%的可复现性相关突变,提出可复现性导向的突变测试作为补充评估方式。
中文摘要 AI 辅助
机器学习研究的可复现性依赖于随机种子、依赖项版本、数据划分和评估配置等实验选择。现有的仓库验证工作流可能成功执行,却未检测到这些选择的变化。我们使用MLReproMutate(一种对机器学习研究仓库应用与可复现性相关的受控突变并对照仓库中现有验证工作流评估的研究软件)研究该问题。我们对39个冻结的仓库操作案例开展了结果盲法实证研究,使用四类突变:随机种子、依赖项固定、数据划分和交叉验证折数。在观察突变结果前,仓库修订、突变候选和验证工作流均已确定。初始执行产生了39个案例中13个的结果;有界恢复程序将可评估的组合集增加到24个。排除1个已确认的等价突变后,剩余23个已确认的非等价突变。选定的验证工作流检测到这23个突变中的2个,对应观察到的检测比例为8.7%。这些结果并不意味着对应仓库不可复现,而是表明在该样本中,现有验证工作流常未检测到本研究引入的特定受控可复现性相关变化。该发现推动了面向可复现性的突变测试,作为评估研究软件保障是否约束实验重要选择的补充方式。
英文摘要
Reproducibility in machine-learning research depends on experimental choices such as random seeds, dependency versions, data partitioning, and evaluation configuration. Existing repository validation workflows may execute successfully without detecting changes to such choices. We study this problem using MLReproMutate, research software that applies controlled, reproducibility-relevant mutations to ML research repositories and evaluates them against validation workflows already present in those repositories. We conducted an outcome-blind empirical study of 39 frozen repository-operator cases using four mutation classes: random seed, dependency pin, data split, and cross-validation fold count. Repository revisions, mutation candidates, and validation workflows were fixed before mutation outcomes were observed. Primary execution yielded outcomes for 13 of 39 cases; a bounded restoration procedure increased the combined evaluable set to 24. After excluding one confirmed-equivalent mutation, 23 confirmed non-equivalent mutations remained. The selected validation workflows detected 2 of these 23 mutations, corresponding to an observed detection proportion of 8.7%. These results do not imply that the corresponding repositories are irreproducible. Rather, they show that, in this sample, existing validation workflows often did not detect the particular controlled reproducibility-relevant changes introduced by the study. The findings motivate reproducibility-oriented mutation testing as a complementary way to assess whether research-software safeguards constrain experimentally important choices.