AI 中文总结
本研究实证探究训练随机性对Rashomon效应的影响,发现不同随机性来源对解空间、预测和解释多样性影响各异,且集合大小与分歧不一致,强调需区分具体来源以评估模型可靠性。
AI 中文摘要
神经网络训练本质上具有随机性,诸如权重初始化等因素会导致尽管预测性能相当,却产生不同的模型。这一现象通常与Rashomon效应相关联,该效应描述了同一任务存在多个接近最优模型的情况。尽管Rashomon效应已受到越来越多的关注,但不同来源的训练随机性对其表现形式的贡献是相似还是不同,仍不清楚。在本工作中,我们从三个互补维度对Rashomon现象进行了实证研究:解空间多样性、预测多样性和决策依据多样性。这些维度分别通过经验Rashomon集合的大小、预测模糊性以及XAI归因图之间的一致性来量化。通过独立控制三种标准的随机性来源,即权重初始化、小批量数据排序和dropout,我们分离了它们各自对Rashomon现象每个维度的贡献。在表格和图像分类基准上的实验表明,这些来源对三个维度的影响方式不同。特别是,更大的经验Rashomon集合并不必然对应更大的预测分歧或更低的解释一致性,这表明解空间、预测和决策依据多样性捕捉的是Rashomon效应的互补方面,而非可互换的方面。总体而言,我们的结果表明,训练随机性不仅影响预测性能,还影响预测和解释的稳定性,这凸显了在评估神经网络模型的可靠性、可复现性和可解释性时,识别导致Rashomon现象不同表现的具体随机性来源的重要性。
英文摘要
Neural network training is inherently stochastic, with factors such as weight initialization leading to distinct models despite comparable predictive performance. This phenomenon is commonly associated with the Rashomon effect, which describes the existence of multiple near-optimal models for the same task. Although the Rashomon effect has received increasing attention, it remains unclear whether different sources of training stochasticity contribute similarly or differently to its manifestations. In this work, we present an empirical study of the Rashomon phenomenon along three complementary dimensions: solution-space multiplicity, predictive multiplicity, and decision-basis multiplicity. These dimensions are quantified through the size of the empirical Rashomon set, predictive ambiguity, and agreement between XAI attribution maps, respectively. By independently controlling three standard sources of stochasticity, namely weight initialization, mini-batch data ordering, and dropout, we isolate their respective contributions to each dimension of the Rashomon phenomenon. Experiments on tabular and image classification benchmarks reveal that these sources affect the three dimensions in different ways. In particular, larger empirical Rashomon sets do not necessarily correspond to greater predictive disagreement or lower explanation agreement, indicating that solution-space, predictive, and decision-basis multiplicity capture complementary rather than interchangeable aspects of the Rashomon effect. Overall, our results show that training stochasticity influences not only predictive performance but also the stability of predictions and explanations, highlighting the importance of identifying the specific sources of stochasticity responsible for different manifestations of the Rashomon phenomenon when assessing the reliability, reproducibility, and interpretability of neural network models.
CommentsSubmitted to a journal for peer review