arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27964cs.LG

线性RNN缩放定律:当更长序列胜过更多序列时

Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences

Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过教师-学生线性RNN模型,推导出序列预训练中的显式缩放定律,揭示序列长度与序列数量在双尺度机制下不可互换,并给出理论证明。

中文摘要 AI 辅助

自回归语言模型的经验缩放定律将预测损失与模型大小、数据大小和优化计算联系起来,但其在序列预训练设置中的理论起源仍知之甚少。我们在一个易于处理的教师-学生模型中研究此问题,其中稳定的潜在线性RNN生成轨迹,而一个草绘线性循环学生通过受保护的全批次WSD梯度下降进行下一词预测训练。草绘维度$M$扮演模型大小的角色,而$N$条长度为$P$的独立轨迹提供训练令牌。我们允许创新和初始化协方差具有不同的幂律指数$α$和$θ$。诱导的设计谱产生由谱交叉分离的显式近似、优化和统计缩放定律。当$θ\geα$时,恢复原始单尺度速率$M^{1-β_α}$、$R^{(1-β_α)/α}$和$(NP)^{-1}\min\{M,R^{1/α}\}$。当$α-2r\leθ<α$时,更重的初始化尾部改变超出$P$依赖的模型和优化交叉的速率。证明仅在内部使用协方差事件,并在其补集上使用全局受保护的步长。方差保留因子$(NP)^{-1}$,而序列长度也抑制初始化瞬态,因此在双尺度机制中$N$和$P$不再完全可互换。

英文摘要

Empirical scaling laws for autoregressive language models relate prediction loss to model size, data size, and optimization compute, but their theoretical origin is still poorly understood in sequential pretraining settings. We study this question in a tractable teacher--student model where a stable latent linear RNN generates trajectories and a sketched linear recurrent student is trained by safeguarded full-batch WSD gradient descent on next-token prediction. The sketch dimension $M$ plays the role of model size, while $N$ independent trajectories of length $P$ provide the training tokens. We allow the innovation and initialization covariances to have different power-law exponents $α$ and $θ$. The induced design spectrum produces explicit approximation, optimization, and statistical scaling laws separated by spectral crossovers. When $θ\geα$, the original one-scale rates $M^{1-β_α}$, $R^{(1-β_α)/α}$, and $(NP)^{-1}\min\{M,R^{1/α}\}$ are recovered. When $α-2r\leθ<α$, the heavier initialization tail changes the rates beyond $P$-dependent model and optimization crossovers. The proof uses a covariance event only internally and a globally safeguarded step size on its complement. The variance retains the factor $(NP)^{-1}$, while sequence length also suppresses the initialization transient, so $N$ and $P$ cease to be fully interchangeable in the two-scale regime.

发表机构

  • The University of Sydney(悉尼大学)
  • Together AI
  • The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑