叶值作为坐标:梯度提升集成的精确对比解释
Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles
AI总结:
该研究将梯度提升集成的叶值视为坐标,提出精确对比解释方法,构建的追索方法在表格数据集上精度高,在信用数据集上表现优于基线,且能生成可执行的建议。
AI中文摘要:
梯度提升集成通过对每棵树的一个叶值求和进行预测。将这些值视为坐标而非中间结果,每个实例就成为模型在R^M上线性作用的点:得分是坐标之和。这一视角的微小改变使对比解释变得精确。两个实例的差异是一个向量,在它们共享叶的地方恒为零,因此被拒申请人与被接受申请人之间的差距由少数坐标承载,每个坐标都可追溯到真实树中的真实分裂。无需对特征进行拟合、采样或假设其可加性——可加性已存在于正确的空间中。我们基于该表示构建了一个追索方法,并在五个表格数据集上通过重复交叉验证对其进行评估。该方法的建议能以6.2×10^-15的精度重构模型自身的决策,因此审计人员无需模型即可重新检查算术运算。在信用数据集上,该方法在努力程度与现实性的权衡中处于帕累托非支配地位。当建议被限制为主体实际可以做出的改变——而非年龄、已确定的 delinquency(拖欠)——时,它保留了58%的有效性,而最强基线仅保留41%,这种区别是标准评估无法察觉的,因为标准评估从不询问建议是否可以被执行。
英文摘要:
A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than as intermediate results, and every instance becomes a point in R^M on which the model acts linearly: the score is the sum of the coordinates. This small change of view makes contrastive explanation exact. The difference between two instances is a vector that is identically zero wherever they share a leaf, so the gap between a rejected applicant and an accepted one is carried by a handful of coordinates, each traceable to a real split in a real tree. Nothing is fitted, sampled, or assumed additive in features -- the additivity is already there, in the right space. We build a recourse method on this representation and evaluate it on five tabular datasets under repeated cross-validation. Its recommendation reconstructs the model's own decision to 6.2 x 10^-15, so an auditor can re-check the arithmetic without the model. On the credit datasets it is Pareto-non-dominated on effort against realism. And when recommendations are restricted to changes the subject could actually make -- not their age, not a settled delinquency -- it retains 58% of its validity where the strongest baseline retains 41%, a distinction the standard evaluation cannot see because it never asks whether a recommendation can be carried out.