arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaRoPE:并非所有注意力头都应同等旋转和缩放

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian Li

arXiv 2607.19363首次发表:更新:

AI 中文总结

研究发现Transformer中不同功能的注意力头需不同频率范围和缩放因子,提出AdaRoPE为各头配备可学习参数,实验表明其性能优于现有变体,能更好地进行上下文扩展并保留短上下文性能,凸显在单头级别优化旋转位置嵌入之重要性。

AI 中文摘要

旋转位置嵌入(RoPE)在Transformer中被广泛用于编码位置信息,但标准实现对所有注意力头强制执行统一的频率调度和缩放。通过简化的检索任务和长度泛化场景,我们从经验和理论上表明,具有不同功能角色的头需要不同的频率范围和注意力缩放因子才能有效运行。忽略这种结构会导致嵌入维度利用次优和性能下降,特别是在长上下文设置下。为解决这些限制,我们提出了AdaRoPE,为每个注意力头配备可学习的旋转频率和注意力缩放因子。使用AdaRoPE的预训练语言模型始终优于现有RoPE变体,包括部分RoPE和NoPE基线。对于上下文扩展,我们进一步表明,YaRN等方法中使用的统一频率和注意力缩放是次优的。通过应用特定于头的缩放,AdaRoPE在推断设置和长上下文持续预训练设置中都能实现更好的上下文扩展,同时更好地保留短上下文性能。这些结果突出了在单个注意力头级别优化旋转位置嵌入的重要性。

英文摘要

Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors. Pretrained LLMs with AdaRoPE consistently outperform existing RoPE variants, including partial RoPE and NoPE baselines. For context extension, we further show that uniform frequency and attention scaling, used in methods such as YaRN, are suboptimal. By applying head-specific scaling, AdaRoPE enables better context extension while better preserving short-context performance in both the extrapolation setting and the long-context continued pretraining setting. These results highlight the importance of optimizing rotary position embedding at the level of individual attention heads.

CommentsAccepted at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑