发表机构
Massachusetts Institute of Technology; University of Pennsylvania(麻省理工学院; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究区分了训练动态指数与最优数据指数,证明尽管动态指数因优化器而异,最优数据缩放指数在不同优化器下均收敛至1/3,并揭示其背后的统一动态指数关系。
AI 中文摘要
神经缩放(即损失随训练以幂律下降)是大语言模型的核心特性,最近有研究提出,1/3 的指数源于对峰值分布的学习。该理论描述的是随机梯度下降(SGD),但实际中的模型通常使用自适应优化器进行训练。在此,我们区分了 1/3 理论未加以区分的两个指数:单次运行中损失随训练步数下降的速度,以及最优调参后的损失随数据集大小 D 下降的速度。我们表明,前者(动态指数)是优化器特定的,而后者(最优数据指数)在不同优化器下均收敛到 1/3。在一个在线教师-学生模型中,我们将损失分解为范数增长(径向)和向教师方向的对齐(切向),两者分别以动态指数 α_r 和 α_t 的幂律衰减。在 SGD 下,两者都接近 1/3,因此数据指数在不同学习率下也是 1/3。在 Adam 下,两者分离:α_r ≈ 0.48,而 α_t ≈ 0.08。由于总损失在这两部分平衡时最小化,最优学习率是优化器相关的:对 SGD 与 D 无关,但对 Adam 随 D 下降。然而,当调整到该最优值时,两者的损失都回到 D^{-1/3}。随机动力学分析解释了原因:优化器可以在两个通道之间交换衰减速度,但它们都落在一个单一动态指数关系 2α_r + α_t = 1 上,这固定了最优数据指数为 1/3。在包括 Muon 在内的七个优化器中,测得的指数与该关系一致,且最优损失包络线在它们之间与 D^{-1/3} 吻合。优化器决定了模型每步学习的速度;在最优调参下,它改变的是前因子,而不是损失随每个样本下降的速率。
英文摘要
Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size $D$. We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to $1/3$ across optimizers. In an online teacher-student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents $α_{r}$ and $α_{t}$. Under SGD, both are close to $1/3$, so the data exponent is also $1/3$ across different learning rates. Under Adam the two separate: $α_{r} \simeq 0.48$ but $α_{t} \simeq 0.08$. Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: $D$-independent for SGD but falls with $D$ for Adam. Yet tuned to that optimum, the loss returns to $D^{-1/3}$ for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, $2α_{r}+ α_{t} = 1$, which fixes the optimal data exponent at $1/3$. Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with $D^{-1/3}$ across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.