arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LeRoPE:可学习的RoPE频率提升语言建模

LeRoPE: Learnable RoPE Frequencies Improve Language Modeling

Petros Karypis, Sean O'Brien, Shreyas Kadekodi, Rui Zhu, Julian McAuley

arXiv 2607.10134首次发表:更新:

发表机构

UC San Diego(加利福尼亚大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究改进语言模型中位置编码的方法,提出可学习RoPE频率的LeRoPE方法,通过训练不同规模语言模型验证,该方法在各规模下均优于RoPE和部分RoPE,提升了语言建模性能。

AI 中文摘要

旋转位置编码(RoPE)是现代语言模型中当前最流行的位置编码。RoPE旋转查询和键向量的二维块,作为它们相对位置偏移的函数运行。RoPE中逐位置的旋转速率通常遵循由固定基频超参数指定的几何序列。先前工作通过增加此参数以减缓旋转或仅将RoPE应用于QK维度的一个子集来提高性能。在这项工作中,我们通过为每个频率学习一个标量来修改RoPE,将频率视为可学习参数而非超参数。我们通过从零开始训练一系列从52M到25亿参数的语言模型来验证可学习的RoPE。我们观察并分析了一个高范数、位置性的LeRoPE频段的出现。LeRoPE在所有规模上始终优于RoPE和部分RoPE,在最大规模下,RoPE需要多3.4%的计算量(FLOPs)才能与LeRoPE匹配。

英文摘要

Rotary Positional Encodings (RoPE) are currently the most popular positional encodings used in modern language models. RoPE rotates two-dimensional chunks of query and key vectors, operating as a function of their relative positional offset. The position-wise rates of rotation in RoPE typically follow a geometric sequence specified by a fixed base-frequency hyperparameter. Prior work has improved performance by either increasing this parameter to slow rotation or by applying RoPE to only a subset of QK dimensions. In this work we modify RoPE by learning a scalar per frequency, treating frequencies as learnable parameters rather than hyperparameters. We validate Learned RoPE by training a ladder of language models from scratch, ranging from 52M to 2.5B parameters. We observe and analyze the emergence of a high-norm, positional LeRoPE band. LeRoPE consistently outperforms RoPE and partial RoPE across all scales, with RoPE requiring 3.4% more compute (FLOPs) to match LeRoPE at the largest scale.

Comments27 pages, 10 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑