发表机构
University of California Irvine(加州大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
qshap是支持R与Python的工具,可对GBDT进行R²的夏普利分解,适配xgboost等主流GBDT实现,还为遗忘树提供专用后端以加速计算,用于量化特征对模型性能的贡献。
AI 中文摘要
目前已开发出大量方法用于量化树集成模型中单个预测的特征归因,但许多应用需要特征对整体模型性能贡献的全局度量。尽管可聚合局部归因分数来表征特征重要性,但此类汇总无法直接分解R²等预测性能度量。本文介绍qshap,其支持R和Python两种版本,可对梯度提升决策树(GBDT)进行R²值的夏普利分解,以量化特定特征对模型性能的贡献。通过分解单个观测值的二次损失,qshap提供灵活工具以探究单个特征和观测值的重要性。qshap当前支持广泛使用的GBDT实现,包括xgboost、lightgbm和catboost,通过统一的树表示和高效的C++后端实现;其模块化设计可适配其他基于二叉决策树构建的GBDT实现。此外,本文还为遗忘树(oblivious trees)引入专用后端,利用其对称结构大幅加速计算。
英文摘要
Numerous methods have been developed to quantify feature attributions in individual predictions for tree ensembles. However, many applications require global measures of feature contributions to overall model performance. Although local attribution scores can be aggregated to characterize feature importance, such summaries do not directly decompose measures of predictive performance, such as $R^2$. This article introduces qshap, available in both R and Python, which provides Shapley decomposition of $R^2$ values for gradient-boosted decision trees (GBDTs) to quantify feature-specific contributions to model performance. By decomposing the quadratic loss of individual observations, qshap provides flexible tools to explore the importance of individual features and observations. qshap currently supports widely used GBDT implementations, including xgboost, lightgbm, and catboost, through a unified tree representation and efficient C++ backends. Its modular design can accommodate other GBDT implementations built from binary decision trees. In addition, we introduce a specialized backend for oblivious trees that exploits their symmetric structure to substantially accelerate computation.