arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

局部混合如何在全局NoPE注意力中编码相对位置

How Local Mixing Encodes Relative Position in Global NoPE Attention

Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge

arXiv 2609.38109首次发表:更新:

发表机构

Zyphra Research(Zyphra研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文解释混合模型(含滑动窗口注意力或门控线性注意力)如何在全局NoPE注意力中隐式编码相对位置,通过近因偏差机制实现,并支持长序列外推。

AI 中文摘要

注意力操作在朴素意义上是位置不变的。然而,位置信息对自然语言至关重要,因此在基于Transformer的模型中,人们开发了多种显式位置编码,如旋转位置编码(RoPE)。尽管长期以来人们一直认为显式位置编码是必需的,但最近的研究表明,在全局注意力层中不编码位置(NoPE)的情况下,交错使用局部混合层(如滑动窗口注意力(SWA)和门控线性注意力)的方法在大规模上取得了成功。这种方法如何以及为何有效,目前尚未得到充分理解。在本文中,我们提出了一种解释,说明此类混合模型如何在全局NoPE层中隐式编码位置。在理论和实证证据的支持下,我们的核心论点是,SWA和门控线性注意力在残差流中诱导了一种近因偏差,该偏差传播到全局注意力logits并被其选择。此外,与仅使用全局NoPE注意力的模型(其中位置信息仅源于因果掩码)中发现的隐式位置编码相比,混合模型中的近因偏差可以在长序列中得以保持。除了加深我们对混合模型如何编码位置的理解外,这些发现还可能为如何以能够无限外推到更长序列长度的方式编码位置提供见解。

英文摘要

The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑