发表机构
Kraków University of Economics(克拉科夫经济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文刻画了注意力展开中 Dobrushin 界取等号的条件(相互自支配),并实证发现该条件在自监督模型和语言模型中普遍成立,但在监督模型和表格模型中成立率较低。
AI 中文摘要
每个注意力展开因子的 Dobrushin 系数满足 $\kappa(\frac12(I+A))\le\frac12(1+\kappa(A))$,将这些不等式在各层上相乘可界定整个展开的系数。我们精确刻画了逐层界取等号的条件:等式成立当且仅当达到 $\kappa(A)$ 的某个 token 对是相互自支配的——即两个 token 中的每一个对自身的注意力强度至少与另一个对它的注意力强度相当。该条件远非自动满足:均匀随机随机矩阵仅在 24-30% 的情况下满足该条件。当对每个单独输入的头平均注意力进行测试,并限制在内容 token(图像块、单词或表格特征,排除 cls、register 和 separator token)上时,该条件在 DINOv2(三种模型规模)、RoBERTa 和 DistilBERT 的每一层的几乎所有输入上都成立。在监督模型 DeiT-B 和 ViT-B/16 中,该条件分别对 91% 和 64% 的输入-层对成立,所有失败均发生在深度较晚的层。在 FT-Transformer 上,针对两个标准表格基准进行训练时,该条件仅对 11-44% 的输入-层对成立。特殊 token 几乎解释了 DINOv2 和语言模型中的所有失败:当包含它们时,该条件仅对 82-97% 的输入-层对成立。
英文摘要
The Dobrushin coefficient of each attention-rollout factor satisfies $κ(\frac12(I+A))\le\frac12(1+κ(A))$, and multiplying these inequalities over layers bounds the coefficient of the whole rollout. We characterise exactly when the layerwise bound is tight: equality holds if and only if some token pair attaining $κ(A)$ is mutually self-dominant - each of the two attends to itself at least as strongly as the other attends to it. The condition is far from automatic: uniformly random stochastic matrices satisfy it only 24-30% of the time. When tested on the head-averaged attention of each individual input and restricted to content tokens - image patches, words or tabular features, excluding cls, register and separator tokens - the condition holds for essentially every input at every layer of DINOv2 (three model sizes), RoBERTa and DistilBERT. In the supervised models DeiT-B and ViT-B/16 it holds for 91% and 64% of input-layer pairs respectively, with all failures occurring late in depth. In FT-Transformer trained on two standard tabular benchmarks it holds for only 11-44% of input-layer pairs. The special tokens account for almost all failures in DINOv2 and the language models: when they are included, the condition holds for only 82-97% of input-layer pairs.
Comments19 pages, 4 figures, 4 tables