arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你的注意力指标回答的是哪个问题?注意力行作为组成数据

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

Marios Papamichalis, Regina Ruane

arXiv 2608.14712首次发表:更新:

发表机构

The Wharton School, University of Pennsylvania; Yale University(宾夕法尼亚大学沃顿商学院; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出注意力指标的汇点处理选择会反转结论,提出将注意力行视为组成数据分离汇点与内容项,明确训练中熵崩溃的原因,确定约定安全场景并发布复现代码。

AI 中文摘要

Transformer注意力矩阵的每一行是对token的概率分布,在训练好的模型中,大部分概率集中在单个“汇点(sink)”token上,通常是第一个token。用于比较注意力行的标准工具(余弦相似度、Jensen-Shannon散度、香农熵)因此取决于一个论文很少报告的选择:保留汇点,还是删除汇点并重归一化。这个选择可能会反转结论。在来自5个家族的10个预训练模型上,关于两个注意力头哪个更相似的判断有17%至47%会随该约定翻转,且标准BERT注意力头聚类流程中最显著的结构是该约定的人为产物。原因是单值摘要混合了两个问题:汇点占据多少注意力,以及剩余注意力如何在内容token间分配。将行视为组成数据可将它们精确分离:Aitchison距离正交分解为汇点项和内容项,熵通过精确恒等式分解,内容距离具有Transformer本身拥有的不变性。这种分离在实践中很重要:训练期间测得的大部分熵崩溃是汇点增长导致的,而非注意力锐化(在70M参数时占下降量的30%,1B参数时占95%,1.4B参数时占79%),且使用错误通道剪枝注意力头会使困惑度升高百倍以上。我们确定了每个约定安全的场景,测试了一个冻结的样本外预测器(一次确认、一次弃权(不执行)、一次失败),并发布了可复现所有数值的代码。

英文摘要

Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first. Standard tools for comparing attention rows (cosine similarity, Jensen--Shannon divergence, Shannon entropy) therefore hinge on a choice papers rarely report: keep the sink, or drop it and renormalize. This choice can reverse conclusions. On ten pretrained models from five families, 17--47% of verdicts about which of two heads is more similar flip with the convention, and the most prominent structure in a standard BERT head-clustering pipeline is an artifact of it. The reason is that one-number summaries mix two questions: how much attention the sink takes, and how the rest is divided among the content tokens. Treating rows as compositional data separates them exactly: the Aitchison distance splits orthogonally into a sink term and a content term, entropy splits by an exact identity, and the content distance is characterized by invariances the transformer itself possesses. The separation matters in practice: most measured entropy collapse during training is the sink growing, not attention sharpening (30% of the drop at 70M parameters, 95% at 1B, 79% at 1.4B), and pruning heads with the wrong channel can inflate perplexity more than a hundredfold. We map where each convention is safe, test a frozen out-of-sample predictor (one confirmation, one abstention, one failure), and release code regenerating every number.

CommentsPreprint under submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑