AI 中文总结
REALM通过蒸馏深度网络为浅层网络并聚类激活模式定义机制,结合线性模型与解释门,实现可解释且稳定的预测,在表格和图像数据上性能具竞争力。
AI 中文摘要
深度ReLU网络是分段仿射映射,将输入空间划分为多个单元,每个单元具有不同的激活模式。这种结构促使在每个单元内拟合局部线性模型,以在保持预测精度的同时提高可解释性。挑战在于识别稳定、数据自适应且易于解释的机制。我们提出REALM,一种线性模型混合体,其机制由神经激活模式诱导。由于深度神经网络(DNN)中的激活单元数量可能随深度快速增长,我们首先将深度教师网络蒸馏为宽而浅的学生网络(WSSN),然后对其隐藏层激活进行二值化和聚类以定义机制,并在每个机制内拟合线性模型。由于机制是从内部结构发现的,路由器不承担预测负担。为使机制分配可解释,我们训练一个多类逻辑回归(解释门)来重现机制分配。这种两级结构在原始表格特征或学习到的卷积特征层面均可解释:门识别决定机制分配的特征,而线性模型识别每个机制内驱动预测的特征。我们分析了一个理想化设置,展示了分区复杂性与稳定性之间的权衡:随着机制数量增加,更细的分区可改善近似,但可能降低机制分配的稳定性。在表格和图像数据集上的实验表明,REALM在与其他DNN引导的混合替代模型和固有可解释模型相比时,实现了具有竞争力的预测性能,同时产生稳定的机制级解释。
英文摘要
Deep ReLU networks are piecewise-affine mappings that partition the input space into cells, each characterized by a distinct activation pattern. This structure motivates fitting a local linear model within each cell to preserve predictive accuracy while improving interpretability. The challenge is to identify regimes that are stable, data-adaptive, and easy to explain. We propose REALM, a mixture of linear models whose regimes are induced by neural activation patterns. Because the number of activation cells in a deep neural network (DNN) can grow rapidly with depth, we first distill a deep teacher into a wide, shallow student network (WSSN), then binarize and cluster its hidden-layer activations to define the regimes and fit a linear model within each regime. Since the regimes are discovered from internal structure, the router does not carry the predictive burden. To make regime assignment interpretable, we train a multiclass logistic regression, the explanatory gate, to reproduce the regime assignments. The two-level structure is interpretable at both stages in terms of raw tabular or learned convolutional features: the gate identifies features that determine regime assignments, while the linear models identify features that drive predictions within each regime. We analyze an idealized setting that illustrates a trade-off between partition complexity and stability: as the number of regimes grows, finer partitions can improve approximation but may reduce regime-assignment stability. Experiments on tabular and image datasets show that REALM achieves competitive predictive performance relative to other DNN-guided mixture surrogates and inherently interpretable models while producing stable regime-level explanations.
Comments18 pages, 12 figures, 3 tables