arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多头自注意力是一种参数识别机制

Multi-Head Self Attention is a Parameter Identification Mechanism

W. Ross Morrow

arXiv 2609.01231首次发表:更新:

AI 中文总结

该研究证明多头缩放点积注意力是参数识别策略,头数越多模型越易识别,还探讨了RoPE、GQA等Transformer改进对参数比值的影响,为Transformer架构提升性能提供了统计解释。

AI 中文摘要

我们证明了,多头缩放点积注意力可被视为一种参数识别策略。未识别参数的数量与总参数数量的比值,与头数的倒数成正比,即从1/2变为1/(2H),这意味着具有更多头的模型在结构上更具可识别性。一个微妙的数学观察结果是,注意力永远无法被完全识别,这是一个副作用。类似地,我们还表明,在单头和多头设置中,一些偏置项对基于softmax的注意力层没有影响,不过这大多是一种值得关注的现象,对模型规模以及模型训练/预测效率的影响应是边际性的。我们还从该视角探讨了Transformer的现代改进,包括RoPE和GQA,说明它们如何提高“有意义”参数与所有参数的比值。简单的数值示例表明,训练确实可能涉及与因缺乏识别而产生的模型不变子空间重叠的更新。作为实验的一部分,我们使用了一种“重新平衡”方法,该方法可“修复”与未识别子空间重叠的更新,但并不试图提供应实际采用该方法的证据。相反,我们仅将数值结果视为对理论结果的探索与验证。总体而言,我们从纯数学/统计解释——识别——的角度,探讨了Transformer中的特定架构选择为何可能提升性能。

英文摘要

We prove that a multi-head scaled dot product attention can be viewed as a parameter identification strategy. The ratio of unidentified parameters to the total number of parameters scales like the reciprocal of the number of heads ($1/2 \to 1/(2H)$), meaning models with more heads are structurally more identified. A subtle side effect of the mathematics observation that attention can never be fully identified. Similarly we also show that some bias terms can have no effect on softmax-based attention layers in both the single- and multiple-head settings, though this is mostly a curiosity that should have a marginal effect on model size and model training/prediction efficiency. We also touch on modern improvements to transformers including RoPE and GQA from this perspective, illustrating how those as well can improve the ratio of ``meaningful'' parameters to all parameters. Simple numerical examples demonstrate that training can indeed involve updates that overlap model-invariant subspaces that arise from a lack of identification. As part of our experiments we use a ``rebalancing'' approach that can ``fix'' updates that overlap unindentified subspaces but do not try to present evidence this should actually be adopted. Instead we simply view our numerical results as exploring and confirming the theoretical results. As a whole we discuss a purely mathematical/statistical explanation, identification, for why specific architectural choices in transformers may have improved performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑