arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LandscapeSHAP:哪个持久同调类获得功劳?

LandscapeSHAP: Which Persistent Homology Class Gets the Credit?

Nikola Milićević

arXiv 2609.31469首次发表:更新:

AI 中文总结

针对拓扑数据分析特征的机器学习模型,提出LandscapeSHAP方法,将模型预测公平归因于持久性图点,线性模型有闭式解,非线性模型用蒙特卡洛采样近似,并证明公理唯一性与稳定性。

AI 中文摘要

Shapley值,作为合作博弈论中的一个解概念,最近已成为机器学习中特征功劳分配的标准工具。它们提供了一种基于公理的方法,在数据特征之间公平地分配模型的预测。Shapley值尚未被应用于解释基于拓扑数据分析特征训练的机器学习模型。我们开发了我们认为的首个此类方法,专注于持久性图的持久性景观特征化。由于每个景观坐标是一个秩统计量,将模型的预测功劳归因于单个持久同调类(持久性图点)并非易事。我们引入了LandscapeSHAP,一种基于模型预测对持久性图点进行公平功劳分配的方法。对于持久性景观上的线性模型,LandscapeSHAP具有闭式表达式,可给出每个持久性图点的精确Shapley值。特别是,无需进行联盟采样。我们进一步证明了四个Shapley“公平性”公理唯一地刻画了这种功劳分配,适用于任何模型,而不仅仅是线性模型。对于一般的非线性模型,这个唯一值只能通过其定义的联盟平均公式精确计算,这需要考虑所有$2^N$个联盟,其中$N$是持久性图中的点数。这对于实际规模的持久性图在计算上是不可行的。我们用高效的蒙特卡洛采样持久性图联盟来补充精确的线性模型结果。我们给出了收敛速率,即逼近到所需精度所需的样本数量。我们还证明了LandscapeSHAP功劳分配对于任何模型的稳定性结果。

英文摘要

Shapley values, a solution concept from cooperative game theory, have recently become a standard tool for feature credit allocation in machine learning. They provide an axiomatically justified method to fairly distribute a model's prediction among the data features. Shapley values have not yet been applied to explain machine learning models trained on features from topological data analysis. We develop what we believe is the first such approach, focusing on the persistence landscape featurization of persistence diagrams. Because each landscape coordinate is a rank statistic, crediting a model's prediction back to individual persistent homology classes (persistence diagram points) is nontrivial. We introduce LandscapeSHAP, a method for fair credit allocation to persistence diagram points based on a model's prediction. For linear models on persistence landscapes, LandscapeSHAP has a closed form expression that gives the exact Shapley value of every persistence diagram point. In particular, there is no coalition sampling required. We further prove that the four Shapley "fairness" axioms uniquely characterize this credit allocation for any model, not only linear ones. For a general nonlinear model, this unique value can only be calculated exactly from its defining coalition averaging formula, which requires considering all $2^N$ many coalitions, where $N$ is the number of points in the persistence diagram. This is computationally intractable for persistence diagrams of realistic size. We complement the exact linear model result with an efficient Monte Carlo sampling of persistence diagram coalitions. We give convergence rates in terms of number of samples needed to approximate to a desired degree of accuracy. We also prove stability results for the LandscapeSHAP credit allocation, for any model.

Comments38 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑