arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

涌现的三分之一幂律缩放:当注意力试图集中时

Emergent One-Third Scaling Law as Attention Tries to Concentrate

Yizhou Liu, Sara Kangaslahti, Jeff Gore

arXiv 2609.32100首次发表:更新:

AI 中文总结

研究发现,LLM中softmax学习峰值分布时,logit幅度以1/3指数幂律增长,成为训练瓶颈,导致整体损失遵循1/3缩放,且注意力头是驱动该缩放的关键。

AI 中文摘要

神经缩放定律将更长的训练与更好的性能通过幂律联系起来,这是当今大型语言模型(LLM)的核心,但其起源仍存在争议。最近的一个提议是,幂律可以从单个softmax头学习峰值分布时的强非线性中涌现。在LLM中,多个softmax函数的情况尚不清楚。在这里,我们通过玩具模型表明,任何学习峰值分布的softmax,无论其在模型中的位置如何,其logit幅度都可以以指数为1/3的幂律增长,成为训练瓶颈,其损失贡献以相同指数1/3的幂律衰减。因此,只要至少有一个softmax学习峰值分布,整体损失就遵循1/3缩放。我们确认LLM中的许多softmax函数学习峰值分布,且LLM损失缩放符合这一1/3预测。此外,logit增长动力学揭示,注意力头而非语言建模头,可能是驱动LLM中1/3损失缩放的瓶颈。注意力试图集中在特定信息上,这是Transformer的核心,因此也可能成为训练神经缩放定律的核心。

英文摘要

The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distributions. What happens with multiple softmax functions, as in LLMs, is unclear. Here, we show through toy models that any softmax learning peaked distributions, regardless of its position in the model, can develop logit magnitudes that grow in a power law with exponent $1/3$, becoming a training bottleneck whose loss contribution decays as a power law with the same exponent $1/3$. The overall loss therefore obeys $1/3$ scaling whenever at least one softmax learns peaked distributions. We confirm that many softmax functions in LLMs learn peaked distributions and that LLM loss scaling matches this $1/3$ prediction. Moreover, logit growth dynamics reveal that attention heads, rather than the language modeling head, are the bottleneck likely driving the $1/3$ loss scaling in LLMs. Attention trying to concentrate on specific information, which is the heart of Transformers, may therefore also be the heart of the neural scaling law of training.

Comments32 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑