arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10420cs.SEcs.LG

模型不稳定性只是需要容忍的噪声还是可以管理的属性?

Is Model Instability just Noise to be Tolerated or a Property that can be Managed?

Amirali Rayegan, Lunxiao Li, Tim Menzies

首次发表
浏览论文内容

中文总结 AI 辅助

研究软件分析中模型不稳定性问题,通过调整标签使用、模型复杂度等方法管理不稳定性,使模型达成一致频率提高,误差标准差下降,推荐质量提升,为减少SBSE不稳定性提供基线并建议将不稳定性作为评估轴。

中文摘要 AI 辅助

在软件分析中,对同一分析进行两次重新运行通常会产生不同的模型和结论。这降低了对模型的信任并限制了其使用。我们发现模型不稳定性是一个主要问题。在127个多目标SE优化问题(12700个测试用例)中,即使在改进设置下,最先进的优化器重复运行也仅在13.7%的测试用例上达成一致。我们认为这种不稳定性不仅是要容忍的噪声,而且是一种可以测量和管理的属性。通过调整标签的使用方式、模型的复杂程度以及分割的评分方式,我们得到的模型达成一致的频率是默认配置的4.8倍。优化误差的标准差平均下降22%(从平均17.4降至13.6),而推荐质量提高而非降低。在质量方面,与默认设置的74个相比,优化后的设置在127个数据集中的119个上在统计上排名靠前。我们还测试了因果和数据局部性干预,发现它们仅提供部分帮助,表明存在剩余稳定性下限。我们的证据表明数据本身(噪声、稀缺标签、代理目标以及数据集允许的许多近似等效模型)对稳定性存在基本限制。我们得出结论,不稳定性应作为SE优化中的标准评估轴,应常规测量、与性能一起报告,并用于校准对任何单次运行的信任。本文中的方法提供了一个基线,可据此判断未来减少SBSE不稳定性的努力。为支持开放科学,我们提供了以下再现包:此https URL

英文摘要

In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability

发表机构

  • North Carolina State University(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑