arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoPE已死,RoPE万岁:迈向可扩展的数据感知位置编码

RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings

Jarod Lévy, Mathurin Videau, Jad Yehya, Jean-Rémi King, Stéphane d'Ascoli, Thomas Moreau

arXiv 2609.34556首次发表:更新:

发表机构

Meta AI; Inria, Université Paris-Saclay(Meta AI; 法国国家信息与自动化研究所,巴黎-萨克雷大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RoPE对邻近令牌的偏见及推断时未见角度问题,提出数据感知RoPE(DaRoPE),在快频带保留RoPE、慢频带用学习坐标,跨多领域实验显示其缓解近期偏差并保持最优或持平性能。

AI 中文摘要

Transformer处理令牌时本身不具备任何顺序概念,这使得位置编码成为基本需求而非架构上的改进。旋转位置编码(RoPE)已成为现代语言模型中默认的位置编码方式,然而它严重偏向于邻近令牌。现有替代方案在不同设置下进行了评估,导致文献碎片化且缺乏明确的替代方案。我们通过考察RoPE的一个特定弱点来为这一领域带来结构:其慢频带,这些频带的波长超过训练上下文,在推断期间使模型暴露于未见角度。因此,我们引入了数据感知RoPE(DaRoPE),它在快频带上保留标准RoPE,但在慢频带上用从上下文表示中学习到的有界坐标替代绝对位置。因此,慢频带的几何形状依赖于数据而非仅依赖于位置距离。我们在匹配条件下,跨合成任务、符号音乐、基因组学、神经信号以及涵盖124M到50B参数的语言模型上比较了代表性编码。在这些实验中,DaRoPE在非文本基准上领先,缓解了近期偏差,同时在语言建模和长度推断方面保持最佳或持平。此外,学习到的坐标也使机制可解释,揭示了注意力层如何利用超出令牌距离的上下文信息。综合这些结果,支持DaRoPE在评估方法中作为最佳整体默认选择,当没有特定领域原因偏好其他方法时。

英文摘要

Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing alternatives have been evaluated under different settings, leaving the literature fragmented and without a clear replacement. We bring structure to this landscape by examining a specific weakness of RoPE: its slow frequency bands, whose wavelengths exceed the training context and expose models to unseen angles during extrapolation. We therefore introduce Data aware RoPE (DaRoPE), which preserves standard RoPE on the fast bands but replaces absolute position on the slow bands with bounded coordinates learned from contextual representations. Therefore, the slow-band geometry depends on the data rather than only on positional distance. We compare representative encodings under matched conditions across synthetic tasks, symbolic music, genomics, neural signals, and language models spanning 124M to 50B parameters. Across these experiments, DaRoPE leads on non-text benchmarks, mitigates recency bias, while remaining best or on par in language modeling and length extrapolation. Moreover, the learned coordinates also make the mechanism interpretable, revealing how attention layers leverage contextual information beyond token distance. Together, these results support DaRoPE as the best overall default among the evaluated methods, when there is no domain-specific reasons to prefer another.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑