arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过集成边距和局部预测变异性衡量一致性:存在预测多样性时的决策系统审计

Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity

Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate

arXiv 2609.01397首次发表:更新:

发表机构

Rutgers University; Pacific Northwest National Lab(罗格斯大学; 太平洋西北国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出结合集成边距与局部预测变异性的一致性准则,用于审计存在预测多样性的决策系统,实验表明其可降低错误预测未被检查的风险,且能可靠捕捉Rashomon集的多样性。

AI 中文摘要

Rashomon效应是一种机器学习现象,指同等准确的模型对相同输入会产生不同预测(即预测多样性)。现有工作主要关注单个模型内部的多样性,但在更复杂的决策系统中,Rashomon效应的影响尚未得到充分理解。本研究从审计错误集成预测的角度探讨多样性,其中将实例分流供人工审核的决策基于一致性准则,该准则结合了集成边距与每个组成模型的局部预测变异性度量。在关于稳定性和平滑性的温和假设下,我们证明随着集成规模和用于测量局部预测变异性的样本数量增加,有限集成的一致性得分会收敛到Rashomon集中期望模型对应的一致性得分。为验证所提准则的有效性,我们针对应用于自然语言理解任务的transformer模型,以及用于表格数据分类任务的大语言模型的参数高效微调,评估了该框架。实验表明,与审计单个模型相比,对来自Rashomon集的模型进行集成可大幅降低错误预测未被检查的风险,同时仅导致分流数量适度增加。此外,完整Rashomon集的审计行为可通过规模相对适中的有限集成紧密近似,部分数据集的风险趋近于零。我们进一步证明,与现有一致性度量相比,所提度量与已确立的预测多样性指标具有更强的一致性,为捕捉Rashomon集中的多样性提供了更可靠的方法。

英文摘要

The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). Existing work primarily focuses on multiplicity within individual models, but in more complex decision systems, the impact of the Rashomon effect is less well understood. In this work, we study multiplicity from the perspective of auditing incorrect ensemble predictions, where the decision to divert an instance for human review is based on a consistency criterion that combines the ensemble margin with a measure of local prediction variability for each constituent model. With mild assumptions about stability and smoothness, we show that the consistency scores of finite ensembles converge to the corresponding consistency score of the expected model from the Rashomon set as the ensemble size and the number of samples used to measure local prediction variability increase. To demonstrate the efficacy of the proposed criterion, we evaluate the framework with respect to transformer models applied to natural language understanding tasks and parameter-efficient fine-tuning of large language models used for tabular data classification tasks. Our experiments show that ensembling models from the Rashomon set substantially reduces the risk of incorrect predictions going unchecked compared with auditing a single model, while incurring only a moderate increase in the number of diversions. Moreover, the auditing behavior of the full Rashomon set can be closely approximated by finite ensembles of relatively modest size, with the risk approaching zero for some datasets. We further demonstrate that the proposed measure exhibits stronger agreement with established predictive multiplicity metrics than existing consistency measures, providing a more reliable way to capture multiplicity in the Rashomon set.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑