RelShap:关系一致的Shapley解释
RelShap: Relationally Consistent Shapley Explanations
浏览论文内容
中文总结 AI 辅助
RelShap是融入关系约束的Shapley解释框架,可提升解释的忠实度,在受控设置中能正确识别主导特征,且运行时可通过等价类优化。
中文摘要 AI 辅助
机器学习流水线通常会将关系型数据展平为单表表示,从而丢弃结构约束。广泛使用的基于Shapley值的特征归因方法依赖于特征独立性,对底层数据中永远不会出现的特征组合评估模型,产生误导性的解释。我们提出RelShap,这是一个将关系约束和数据来源融入Shapley值计算的框架,将背景数据和联盟评估都限制在关系上有效的配置中。该框架与估计器无关,可与Kernel SHAP、Monte Carlo和Leverage SHAP组合使用,且不改变它们的采样或加权属性。函数依赖进一步在特征联盟上诱导等价类,RelShap利用这些等价类在不改变Shapley值的情况下减少运行时间;我们提供了预期加速比的组合特征描述。在多个数据集、模型和估计器上的实验表明,RelShap生成的解释更符合数据生成过程,在受控设置中能正确识别主导特征,而包括Conditional SHAP和ManifoldShap在内的现有方法则无法做到这一点。我们的代码可在:this https URL获取。
英文摘要
Machine learning pipelines commonly flatten relational data into single-table representations, discarding structural constraints. Widely used Shapley value-based feature attributions then rely on feature independence, evaluating the model on combinations that could never arise in the underlying data, producing misleading explanations. We propose RelShap, a framework that incorporates relational constraints and data provenance into Shapley value computation, restricting both background data and coalition evaluation to relationally valid configurations. The framework is estimator-agnostic and composes with Kernel SHAP, Monte Carlo, and Leverage SHAP without altering their sampling or weighting properties. Functional dependencies further induce equivalence classes over feature coalitions, which RelShap exploits to reduce runtime without changing Shapley values; we provide a combinatorial characterization of the expected speedup. Experiments across multiple datasets, models, and estimators show that RelShap produces explanations that are more faithful to the data-generating process, correctly identifying the dominant feature in controlled settings where existing methods, including Conditional SHAP and ManifoldShap, do not. Our code is available at: https://github.com/duneag2/relshap.