MCIR:一种具有可靠性保证的基于特征依赖的可解释性方法
MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees
- University of Oslo(奥斯陆大学)
- Aalborg University(奥尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MCIR-M提出一种依赖感知的全局特征重要性方法,通过条件化强依赖邻居计算归一化比率,在多重共线性下提供稳定排序,并在合成与真实基准上验证了其有效性。
AI中文摘要:
现代机器学习模型通常包含强依赖或冗余的特征,这使得特征归因变得困难,因为共享的预测信息可能分布在相关预测变量之间。现有的方法如SHAP、LIME、HSIC、MI/CMI和SAGE,在多重共线性或近似重复预测变量的情况下,可能产生不稳定的排序。我们提出了互相关影响比方法(MCIR-M),一种依赖感知的全局特征重要性方法,它量化了每个特征在选定依赖邻域之外贡献的独特预测信息。MCIR-M引入了互相关影响比(MCIR),该方法将每个特征条件于强依赖的邻居,并计算条件信息与块级信息的归一化比率。总体得分位于[0,1]区间,在精确条件冗余下等于零。我们还引入了一种轻量级估计程序,该程序使用可用数据的一小部分来计算MCIR,并评估与全数据解释的一致性。在受控的合成冗余实验和UCI HAR基准测试中,MCIR表现出依赖感知的排序行为,其在注入近似重复预测变量的情况下优势最为明显。与独立和条件SHAP、SAGE、HSIC、基于MI的得分以及CIR系列基线的比较在真实数据标准上表现不一。减少解释样本降低了所评估配置中的计算负担,而与全数据解释的一致性则通过排序、头部集和忠实度诊断分别评估。总体而言,MCIR-M为强特征依赖下的全局解释提供了一种实用的依赖感知诊断方法。
英文摘要:
Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.