多站点数据块缺失特征下的Shapley值估计
Shapley Value Estimation for Multi-Site Data with Blockwise-Missing Features
浏览论文内容
中文总结 AI 辅助
针对多站点数据块缺失特征问题,提出FUSHAP方法,利用部分观测辅助站点降低Shapley值估计方差,无需插补,并通过筛选排除不兼容站点,显著降低MSE。
中文摘要 AI 辅助
基于Shapley值(SV)的方法是机器学习中特征归因的主流框架,然而现有的群体级Shapley估计器通常假设用于评估联盟博弈的观测在共同特征空间下是完全观测的。这一假设在生物医学、社会科学和环境监测的多站点研究中经常被违反,因为这些研究中的机构在不同的协议下记录不同的特征,从而在数据源之间产生系统性的块状缺失。我们首先表明,在计算Shapley值之前对缺失特征进行插补的标准补救方法会给最终的归因引入系统性的、依赖于联盟的偏差。然后,我们提出FUSHAP(从部分观测数据中进行融合Shapley归因),该方法利用部分观测的辅助站点来降低初步单站点Shapley估计的方差,而无需插补。基于排列的筛选步骤检测并排除数据分布与目标群体不兼容的站点。在合成实验中,FUSHAP的MSE比单站点估计器低3至8倍,比插补基线低2至3倍,且不引入插补引起的偏差;筛选程序在中等错位下以82%的效能识别错位站点,在强错位下以100%的效能识别。在多站点空气质量和多中心临床数据上,FUSHAP相对于单站点估计器将MSE降低了约3至7倍;在临床应用中,标准插补可能使MSE高于单站点基线。
英文摘要
Shapley value (SV)-based methods are the prevailing framework for feature attribution in machine learning, yet existing population-level Shapley estimators generally assume that observations used to evaluate the coalitional game are fully observed under a common feature space. This assumption is routinely violated in multi-site studies across biomedicine, social science, and environmental monitoring, where institutions record different features under different protocols, producing systematic blockwise missingness across sources. We first show that the standard remedy of imputing missing features before computing Shapley values introduces systematic, coalition-dependent bias into the resulting attributions. We then propose \textbf{FUSHAP} (\textbf{Fu}sion \textbf{Sh}apley \textbf{A}ttribution from \textbf{P}artially-observed data), a method that leverages partially-observed auxiliary sites to reduce the variance of a preliminary single-site Shapley estimate without imputation. A permutation-based screening step detects and excludes sites whose data distributions are incompatible with the target population. In synthetic experiments, FUSHAP achieves $3$--$8\times$ lower MSE than the single-site estimator and $2$--$3\times$ lower MSE than imputation baselines without incurring imputation-induced bias, and the screening procedure identifies misaligned sites with $82\%$ power at moderate misalignment and $100\%$ for strong misalignment. On multi-site air quality and multi-center clinical data, FUSHAP reduces MSE by approximately $3$--$7\times$ relative to the single-site estimator; in the clinical application, standard imputation can increase MSE above the single-site baseline.
发表机构
- Centre for Biomedical Data Science, Duke-NUS Medical School(杜克-新加坡国立大学医学院生物医学数据科学中心)
- School of Data Science, Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院)
- Columbia University(哥伦比亚大学)
- National University of Singapore(新加坡国立大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。