同样优秀,却各不相同:AutoML 包中Rashomon集的基准测试
Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages
浏览论文内容
中文总结 AI 辅助
针对AutoML中x-hacking风险,提出ARSA ML框架量化Rashomon集结构,基准测试发现AutoGluon与H2O存在结构不对称,H2O用户更易受x-hacking影响。
中文摘要 AI 辅助
Rashomon效应描述了存在多个接近最优的模型,这些模型在实现相当性能的同时,却提供根本不同的解释。这在AutoML中造成了一个关键漏洞:x-hacking,即基于模型的解释而非预测价值,选择性地事后挑选模型。现有的AutoML框架均未暴露这一风险。我们引入了ARSA ML,一个开源的Python框架,用于量化AutoML流水线中的Rashomon集结构和预测多重性。利用ARSA ML,我们对AutoGluon和H2O在28个二分类数据集上进行了基准测试,并进行了事后x-hacking分析,揭示了一致的结构不对称性:AutoGluon产生更大、更多样化的集合,且解释稳定,而H2O生成的集合紧凑,但预测分歧和解释不稳定性显著更高——这使得H2O用户更容易遭受x-hacking。这一差距在所有评估指标和epsilon阈值中持续存在,指向了每个框架模型构建策略的根本差异。ARSA ML可在https://this URL获取。
英文摘要
The Rashomon effect describes the existence of multiple near-optimal models that achieve comparable performance while offering fundamentally different explanations. This creates a critical vulnerability in AutoML: x-hacking, the selective post-hoc choice of a model based on its explanation rather than predictive merit. No existing AutoML framework exposes this risk. We introduce ARSA ML, an open-source Python framework that quantifies Rashomon set structure and predictive multiplicity within AutoML pipelines. Using ARSA ML, we benchmark AutoGluon and H2O across 28 binary classification datasets, and conduct a post-hoc x-hacking analysis revealing a consistent structural asymmetry: AutoGluon produces larger, diverse sets with stable explanations, while H2O generates compact sets with markedly higher prediction divergence and explanation instability -- making H2O users considerably more exposed to x-hacking. This gap persists across all evaluated metrics and epsilon thresholds, pointing to a fundamental difference in each framework's model-building strategy. ARSA ML is available at https://pypi.org/project/arsa-ml/ .
发表机构
- Warsaw University of Technology(华沙理工大学)
- Systems Research Institute, Polish Academy of Sciences(波兰科学院系统研究所)
- Eskisehir Technical University(埃斯基谢希尔技术大学)
机构由 AI 辅助整理,请以论文原文为准。