解耦旋转位置嵌入(RoPE)的表达能力
Disentangling the Expressivity of RoPE
- Toyota Technological Institute at Chicago(芝加哥丰田技术学院)
- ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究解耦RoPE的两种核心观点,形式化其表达性,发现常规RoPE的非周期性旋转仅能实现有界模拟,而周期性RoPE可泛化模块化语言,为RoPE Transformer提供更贴近实际的理论表征。
AI中文摘要:
在解释旋转位置嵌入(RoPE)成功的研究中,反复出现两种观点:表达性研究将周期性位置信息与模块化谓词关联,而机制与长上下文研究则强调位置锚点和局部偏移。我们针对完全均匀、有限精度的软注意力Transformer形式化了这两种观点。我们发现,若每个旋转组件都是周期性的,RoPE Transformer恰好能识别用带模块化谓词的过去时态逻辑可定义的语言。常规RoPE则不同:其计算的旋转永不重复,这会产生依赖精度的固定偏移回溯算子的有界模拟,而非全长度的模块化表征。控制实验验证了这种分离:构造的周期性序列在模块化语言上具备长度泛化能力,而常规RoPE表现更像有界局部偏置,会损害需要访问长距离上下文的位置不变性任务。总体而言,我们的发现阐明了RoPE Transformer,使理论表达性表征更贴近实际使用的模型。
英文摘要:
Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize positional anchors and local offsets. We formalize both accounts for fully uniform, finite-precision soft-attention transformers. We find that, if every rotary component is periodic, RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates. Conventional RoPE is different: The rotations it computes never repeat. This yields a precision-dependent bounded simulation of fixed-offset look-back operators, rather than an all-length modular characterization. Controlled experiments match this separation: Constructed periodic schedules length-generalize on modular languages, while conventional RoPE behaves more like a bounded locality bias and can impair tasks requiring position-invariant access to distant context. Altogether, our findings shed light on RoPE transformers, bringing theoretical expressivity characterizations closer to models used in practice.