发表机构
Mississippi State University(密西西比州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究Transformer中长度泛化现象,通过优化解释,证明旋转编码使注意力逻辑仅与相对偏移有关,绝对编码则不然,还将相关现象和特性转移到多层多头Transformer,联系了注意力隐式偏差等多方面内容。
AI 中文摘要
具有相对位置编码的Transformer通常能外推到比训练中所见更长的序列,而具有学习到的绝对编码的Transformer通常不行。这是一个稳健的经验规律,目前对此的解释主要是关于表达能力。我们给出一个优化解释。在一个隔离位置选择的最小固定偏移检索任务中,差距由训练的注意力头的隐式偏差决定。我们证明旋转编码使注意力逻辑仅成为相对偏移的函数,是一种精确的等变性。学习到的绝对编码则让超出范围的位置不受约束。我们将学习到的旋转规则表征为与目标偏移对齐的低秩“载体”核,并推导出由此产生的优雅精度衰减作为注意力稀释定律。线性注意力控制表明该机制特定于softmax。该现象、等变性和载体都转移到在全序列长度泛化任务上训练的多层、多头Transformer。这个解释将注意力的隐式偏差、循环模型中外推的隐式偏差以及RASP-L猜想的学习方面联系起来。
英文摘要
Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not. This is a robust empirical regularity, and the explanations offered for it so far are chiefly about expressivity, that is, about whether a length-generalizing solution exists. We give an optimization explanation. On a minimal fixed-offset retrieval task that isolates positional selection, the gap is governed by the implicit bias of the trained attention head: among the many solutions that fit short sequences, which one gradient descent actually selects. We prove that rotary encodings make the attention logit a function of relative offset alone, an exact equivariance, so whatever selection rule is learned at training lengths is reproduced verbatim at every longer length. Learned absolute encodings instead leave out-of-range positions unconstrained, and the trained head pins to a fixed absolute position inside the training range. We characterize the learned rotary rule as a low-rank ``carrier'' kernel aligned with the target offset, and we derive the resulting graceful accuracy decay as an attention-dilution law; both predictions are confirmed across seeds and offsets. A linear-attention control shows the mechanism is specific to softmax: without normalization, training selects a min-norm interpolant that does not extrapolate. The phenomenon, the equivariance, and the carrier all transfer to a multi-layer, multi-head transformer trained on a full-sequence length-generalization task. The account connects the implicit bias of attention, implicit bias for extrapolation in recurrent models, and the learning side of the RASP-L conjecture.