arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多元预测不确定性的熵-沙普利博弈层级

A Hierarchy of Entropy-Shapley Games for Multivariate Predictive Uncertainty

Niklas Koenen, Claudia Battistin, Jeriek Van den Abeele, Martin Jullum

arXiv 2609.35217首次发表:更新:

发表机构

Leibniz Institute for Prevention Research and Epidemiology – BIPS; University of Bremen; Simula Research Laboratory; Telenor Research & Innovation; Norwegian Computing Center(莱布尼茨预防研究与流行病学研究所——BIPS; 不来梅大学; 西穆拉研究实验室; 挪威电信研究与创新部; 挪威计算中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出熵-沙普利博弈层级,用于多元预测不确定性归因,分离跨分量依赖效应,并提供闭式与样本估计器,验证于多种概率模型。

AI 中文摘要

现代概率机器学习模型越来越多地产生具有复杂依赖结构的多元输出,从多步时间序列预测到样本路径预测。理解哪些输入特征驱动预测不确定性,对于风险感知决策、模型诊断以及决定不确定性应被缓解还是对冲至关重要。这一归因问题需要选择如何处理输出分量之间的依赖关系。现有方法在归因前通过聚合或投影将输出简化为标量,从而掩盖了特征是否影响边际不确定性、依赖结构或两者兼有,而逐分量分析可能完全遗漏依赖效应。我们通过引入三个基于熵的沙普利博弈层级来填补这一空白,该层级为任何有序多元结果明确地做出了这种输出侧选择,范围从逐分量边际熵到完全联合熵。该层级分离出一个跨分量归因项,用于捕捉每个特征如何改变输出分量之间的依赖关系,这是逐分量方法无法观测的量。我们建立了联合归因的链式法则分解,并通过条件总相关刻画了跨分量项,提供了闭式解和基于样本的估计器。最后,我们展示了该框架如何捕捉从分布回归到零样本时间序列基础模型等概率模型中学到的联合结构的差异。

英文摘要

Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑