Transformer中的位置编码:从绝对与相对方法到旋转位置嵌入及长上下文扩展
Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling
浏览论文内容
中文总结 AI 辅助
本技术综述梳理 Transformer 各类位置编码方法,推导 RoPE 原理并对比不同方法特性,研究长上下文扩展方案,明确长上下文泛化能力需多任务评估。
中文摘要 AI 辅助
自注意力模型处理 token 间依赖内容的交互,但本身不编码 token 顺序。位置编码通过在 Transformer 的表示和注意力分数中引入绝对坐标、相对距离或位置相关旋转解决该局限。本技术综述对正弦和学习型绝对位置嵌入、Shaw 式相对位置表示、Transformer-XL、T5 相对位置偏置、ALiBi 及旋转位置嵌入(RoPE)进行统一阐述。推导 RoPE 将绝对位置索引转换为 Query-Key 内积中相对相位差的方式,并从位置注入位置、计算成本、与 KV 缓存的兼容性及长度外推性等方面比较上述方法。随后研究长上下文扩展方法,包括位置插值、RoPE 缩放定律、NTK 感知缩放、动态 NTK、分块 NTK、YaRN、LongRoPE 及 LongRoPE2,重点关注频率分配、注意力重缩放、训练长度和目标上下文长度。还总结了代表性大语言模型中位置编码的实现考量、评估协议及位置编码选择。核心结论为:计算超出训练长度的位置特征的能力不代表具备可靠的长上下文泛化能力;上下文扩展需通过短上下文保留、逐位置困惑度、检索、推理及长上下文代码任务进行评估。
英文摘要
Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or position-dependent rotations into Transformer representations and attention scores. This technical survey develops a unified account of sinusoidal and learned absolute position embeddings, Shaw-style relative position representations, Transformer-XL, T5 relative position bias, ALiBi, and Rotary Position Embeddings (RoPE). We derive how RoPE converts absolute position indices into relative phase differences in Query-Key inner products and compare these methods in terms of where position is injected, computational cost, compatibility with KV caching, and length extrapolation. We then examine long-context extensions, including Position Interpolation, RoPE scaling laws, NTK-aware scaling, Dynamic NTK, NTK-by-parts, YaRN, LongRoPE, and LongRoPE2, with emphasis on frequency allocation, attention rescaling, training length, and target context length. We also summarize implementation considerations, evaluation protocols, and position-encoding choices in representative large language models. A central conclusion is that the ability to compute positional features beyond the training length does not imply reliable long-context generalization; context extension must be evaluated through short-context retention, position-wise perplexity, retrieval, reasoning, and long-context code tasks.