发表机构
A Carrot, Inc.(A Carrot公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于内容的单元寻址方法,解决RoPE在长上下文中的位置外推问题,在浅层实验中显著降低困惑度,并能检索多事实信息。
AI 中文摘要
旋转位置编码(RoPE)利用每个词元的整数位置来决定注意力内部施加的旋转。这对于局部词元顺序效果良好,但上下文长度的增加会造成位置上的训练与测试不匹配:RoPE会在训练期间未见过的偏移处产生相对旋转。那些对位置进行缩放、插值、随机化或偏置的方法,规定了注意力如何处理这些偏移,但仍从不断增长的词元计数器中获取位置信息。我们转而将词元流划分为若干单元,在每个单元内部保留普通的RoPE位置,并为每个已完成的单元分配一个由其内容计算出的地址。添加单元时,便对新的内容应用相同的已学习映射,而非扩展位置范围或标识符表。我们证明,这种构造精确保留了局部RoPE,在插入或重排其他单元时,两个固定词元之间的注意力比较保持不变,并且不会仅仅因为添加更多单元而产生新的相对旋转。在字符级Tiny Shakespeare诊断实验中,一个在256字符上下文上训练的模型在256字符处的验证困惑度为4.04,在4096字符处为3.82,而连续RoPE则从4.71变化到12.09。第二个诊断表明,基于内容的寻址能够检索并使用来自多个序列化事实的信息。这些是受控的浅层实验,而非规模基准测试,但它们支持一个直接的处方:在局部使用位置寻址,在单元之间使用内容寻址。
英文摘要
Rotary position embedding (RoPE) uses each token's integer position to determine the rotation applied inside attention. This works well for local token order, but increasing context length creates a positional train-test mismatch: RoPE produces relative rotations at offsets not seen during training. Methods that rescale, interpolate, randomize, or bias positions specify how attention handles those offsets, but still derive positional information from a growing token counter. We instead divide a token stream into units, retain ordinary RoPE positions within each unit, and assign every completed unit an address computed from its content. Adding units then applies the same learned map to new content rather than extending a positional range or an identifier table. We prove that this construction preserves local RoPE exactly, leaves the attention comparison between two fixed tokens unchanged when other units are inserted or reordered, and does not create new relative rotations merely because more units are added. In a character-level Tiny Shakespeare diagnostic, all-token validation perplexity remains approximately constant from contexts of 256 to 4096 characters. A second diagnostic shows that content-based addressing can retrieve and use information from multiple serialized facts. These are controlled shallow experiments, not scale benchmarks, but they support a direct prescription: use position to address locally and content to address across units.