arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22367cs.LGcs.AI

通过注意力展开实现内部可解释性:Transformer中的收缩和传播概况

Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

  • University of Deusto(德乌斯托大学)
  • Universidad Carlos III de Madrid(马德里卡洛斯三世大学)
  • Universidad Autónoma de Madrid(马德里自治大学)

机构由 AI 辅助整理,请以论文原文为准。

Umberto Biccari, Qian Huang, Enrique Zuazua

AI总结:

研究Transformer内部可解释性,引入基于传播视角的内部可解释性,用注意力展开实现。通过收缩理论分析其传播概况,在代谢组年龄预测模型中有应用,还与其他方法比较,揭示了变量一致性情况,用作注意力介导传播诊断。

AI中文摘要:

特征归因方法为输入变量与模型输出分配分数,但本身并未描述明确定义的交互算子如何在中间层组合。我们引入了内部可解释性,这是一种基于传播的内部模型组织视角,并使用注意力展开为表格Transformer实例化。我们将展开解释为编码特征令牌之间注意力介导传播的行随机算子。通过应用经典的多布林-多布鲁申收缩理论,我们表明具有小多布鲁申系数的展开算子在数量上接近秩一随机矩阵,其公共行由其归一化列和确定。该结果为相应的展开传播概况提供了结构解释。在用于代谢组年龄预测训练的Transformer中,测量的展开收缩随深度增强。训练和随机初始化的模型也表现出不同的传播概况,尽管当前实验未确定单个展开排名变量的预测相关性。与PCA和SHAP的GradientExplainer近似的探索性比较揭示了高排名变量之间的局部一致性,但完整排名之间的一致性较弱。因此,注意力展开在这里用作注意力介导传播的诊断,而不是作为完整Transformer的因果解释或忠实归因。

英文摘要:

Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretability}, a propagation-based perspective on internal model organization, and instantiate it for tabular Transformers using attention rollout. We interpret rollout as a row-stochastic operator encoding attention-mediated propagation between feature tokens. By applying classical Doeblin--Dobrushin contraction theory, we show that a rollout operator with a small Dobrushin coefficient is quantitatively close to a rank-one stochastic matrix whose common row is determined by its normalized column sums. This result gives a structural interpretation to the corresponding rollout propagation profile. In Transformers trained for metabolomic age prediction, the measured rollout contraction strengthens with depth. Trained and randomly initialized models also exhibit different propagation profiles, although the present experiments do not establish the predictive relevance of individual rollout-ranked variables. Exploratory comparisons with PCA and GradientExplainer approximations to SHAP reveal localized agreement among highly ranked variables but weak agreement across complete rankings. Attention rollout is therefore used here as a diagnostic of attention-mediated propagation, not as a causal explanation or faithful attribution of the complete Transformer.

↑