机制转变:位置编码选择如何影响上下文内检索
Shifting Mechanisms: How Positional Encoding Choice Shapes In-Context Retrieval
浏览论文内容
中文总结 AI 辅助
本研究通过机制性分析揭示,位置编码选择(如RoPE与混合PE)会改变语言模型上下文内检索的内部机制,从位置检索转向语义检索,并带来多目标检索增益但损害竞争键区分。
中文摘要 AI 辅助
语言模型越来越多地采用跨层改变注意力跨度和位置编码的架构,例如在滑动窗口注意力中应用RoPE,在全局注意力中应用NoPE(SWA NoPE)。然而,这些选择如何影响上下文内检索仍不清楚。为研究此问题,我们采取机制性视角,追踪位置编码(PE)选择如何塑造模型用于上下文内检索的内部机制。在跨越八个家族的22个开放权重模型中,我们发现标准RoPE模型主要依赖位置检索,而PE混合模型转向语义检索。我们进一步在受控预训练消融实验中表明,将位置编码限制在局部层会产生这种语义转变,并削弱位置信息的表征。最后,我们展示了PE混合模型报告的长上下文增益掩盖了一种检索权衡:SWA NoPE在多目标检索和问答上优于RoPE,但在区分竞争键时性能下降。我们表明,这些行为差异更好地追踪了从位置机制向语义机制的机制转变,而非长上下文检索的均匀改进。
英文摘要
Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding-window attention and NoPE with global attention (SWA NoPE). However, how these choices shape in-context retrieval remains unclear. To study this question, we take a mechanistic view, tracing how positional encoding (PE) choice shapes the internal mechanisms models use for in-context retrieval. Across 22 open-weight models spanning eight families, we find that standard RoPE models rely primarily on positional retrieval, while PE hybrids shift toward semantic retrieval. We further show on a controlled pre-training ablation that confining positional encoding to local layers produces this semantic shift, degrading representations of positional information. Finally, we show that the reported long-context gains of PE hybrids mask a retrieval trade-off: SWA NoPE improves over RoPE on multiple-target retrieval and QA, but degrades when distinguishing competing keys. We show that these behavioral differences better track the mechanism shift from positional toward semantic mechanisms than a uniform improvement in long-context retrieval.
发表机构
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。