arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19107cs.LG

模型增长、递归与边界算子如何影响缩放指数

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Zixi Chen, Akshay Vegesna, Samip Dahal, Andrew Gordon Wilson

首次发表
浏览论文内容

中文总结 AI 辅助

该研究证明架构干预(如模型增长、循环和边界算子)可改变预训练缩放指数,带来计算效率的指数级提升,其中模型增长效果最显著,且增益随规模增加。

中文摘要 AI 辅助

缩放定律预测损失如何随计算量的增加而降低。我们表明,与传统观念相反,架构干预可以修改预训练中的缩放指数,从而在计算量增加时带来性能的指数级提升。作为锚定点,我们考虑了循环变压器的架构形式。尽管通常不这样使用,循环(也称为递归深度)通过增加训练期间的循环次数,提供了一种模型增长的机制。无论是否共享权重,模型增长都对缩放指数产生了最大的改变。特别是,一个7.4B参数的模型增长架构在CORE数据集上以约$20\ imes$更少的计算量匹配了GPT-3 13B的性能,并且其计算效率增益随规模增加而增大。此外,仅仅在普通变压器中使用边界算子(该算子归一化并注入早期块)也能提供递增的计算效率增益,尽管程度较小。在数据受限的多轮训练设置中,标准循环具有有用的正则化效果,我们发现随着规模增加,增加循环次数是计算最优的。这些结果可以通过计算深度的视角来理解:对于给定的计算预算,我们希望增加变压器的可用深度,这可以带来随规模增加而增大的效率增益。

英文摘要

Scaling laws predict how loss decreases with increases in computation. We show, contrary to conventional wisdom, that architectural interventions can modify scaling exponents in pre-training, leading to power-law improvements in performance as computation increases. As an anchoring point, we consider the architectural formulation of looped transformers. Although not typically used in this way, looping, also known as recursive depth, provides a mechanism for model growth, by increasing the number of loops during training. Model growth, with and without shared weights, provides the biggest changes to the scaling exponents. In particular, a 7.4B model growth architecture matches GPT-3 13B on CORE with roughly $20\times$ less compute, and has compute efficiency gains that increase with scale. Moreover, simply using a boundary operator in a vanilla transformer, which normalizes and injects an earlier block, also provides an exponent increase, although to a lesser extent. In the data-constrained, multi-epoch setting, standard looping has a useful regularizing effect, where we find it is compute-optimal to increase the number of loops with scale. These results can be understood through the lens of computational depth: for a given computational budget, we wish to increase the usable depth of the transformer, which can lead to efficiency gains that increase with scale.

发表机构

  • Q Labs(Q实验室)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑