arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于局部蒸馏的可解释人工智能

Interpretable AI with Local Distillation

Erin Craig, Yiling Huang, Snigdha Panigrahi

arXiv 2608.23538首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出局部蒸馏方法,以黑箱教师模型指导正则化线性学生模型,在17个基准数据集上近匹配教师模型准确性,还能识别癌症患者基因表达异质性。

AI 中文摘要

表格基础模型和梯度提升集成等现代AI模型的预测能力优于传统方法,但难以解释其预测依据。高风险决策要求模型兼具准确性与可解释性,局部线性建模为此提供了可行路径:平滑回归函数在局部可被线性函数良好近似,因此在每个查询点附近拟合线性模型,既能保证准确性又不损失透明度。其挑战在于学习何为“局部”,并开发用于解释的统计工具。本文提出局部蒸馏方法,其中黑箱“教师”模型在每个查询点指导正则化线性“学生”模型:教师通过对预测结果相似的训练观测值赋予更高权重来定义局部性,并将其在查询点的预测作为伪观测值加入拟合,该伪观测值的权重由数据估计。为实现解释,我们在局部目标中加入少量高斯随机化,并通过重新拟合评估稳定性:选择频率可识别查询点处的可靠特征,对随机化拟合进行聚类可识别数据中的稳定子组。在套索惩罚下,我们证明这种随机化产生的特征选择概率在训练响应的微小扰动下保持稳定。在17个基准数据集上,局部蒸馏方法的准确性几乎与其教师模型相当,同时在每个测试点生成稀疏线性模型。在高维癌症基因表达示例中,该框架识别出局部模型使用不同基因的患者亚组,这种异质性是全局线性模型无法察觉,且在黑箱模型中难以凸显的。

英文摘要

Modern AI models such as tabular foundation models and gradient-boosted ensembles can outpredict classical methods, but provide little basis for reasoning about their predictions. High-stakes decisions call for models that are both accurate and interpretable as built. Local linear modeling offers a path forward: a smooth regression function is locally well approximated by a linear one, allowing a linear fit near each query point to achieve high accuracy without sacrificing transparency. The challenges lie in learning what is "local" and developing statistical tools for interpretation. Here, we propose local distillation, in which a black-box "teacher" guides a regularized linear "student" model at each query point. The teacher (1) defines locality by upweighting training observations with similar predicted outcomes, and (2) anchors the fit with its prediction at the query point, included as a pseudo-observation whose weight is estimated from the data. For interpretation, we add a small amount of Gaussian randomization to the local objective and use refits to assess stability: selection frequencies identify reliable features at a query point, and clustering the randomized fits identifies stable subgroups across the data. Under the lasso penalty, we prove that this randomization yields feature-selection probabilities that are stable under small perturbations of the training responses. Across 17 benchmark datasets, local distillation nearly matches its AI teacher's accuracy while producing a sparse linear model at each test point. In a high-dimensional cancer gene expression example, the framework identifies patient subgroups whose local models use different genes; this heterogeneity is invisible to a global linear model, and difficult to surface in a black-box model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑