小规模实验:我们达到目标了吗?
Small-Scale Experiments: Are We There Yet?
- FAIR at MSL Meta
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究指出小规模模型对超参数的敏感性是缩放定律在小规模实验中失效的关键,开发了以模型为中心的研究方法论,通过小规模实验证实Transformer中预归一化随模型规模增大效果更好,表明小规模实验可兑现缩放定律的承诺。
中文摘要 AI 辅助
缩放定律曾承诺开展高性价比的实验,但六年过去,它们尚未完全兑现这一承诺。相反,研究人员发现缩放定律在小规模模型(参数规模从4M开始)上不可靠,得出结论称无法避免使用大规模模型。我们证明情况并非如此:混淆因素是超参数。小规模模型对超参数高度敏感,但超参数的敏感性会随模型规模增大而减弱。这种小规模敏感性使得缩放定律容易被忽略,因为缩放定律仅在完全调优的前沿区域才会显现,而要达到该前沿区域需要进行远超多数研究者实际运行的广泛搜索。通过对基本缩放定律方案进行消融实验,我们表明调优良好的超参数比其他任何因素都更重要。此外,我们揭示了为何那些超参数更容易找到:随着规模增大,超参数损失曲面的维度会降低。不过,尽管小规模模型中存在缩放定律,但外推会遇到统计局限性,需要采用整体方法。我们将自身见解与近期文献相结合,开发了一种以模型为中心的新研究方法论,并在一个曾耗费该领域数年才解决的问题上对其进行验证:在Transformer架构中应将归一化层放在何处。从小规模实验中,我们得到了大规模实验的结果:随着模型规模增长,预归一化效果更好。借助合适的工具和更深入的理解,小规模实验能够兑现缩放定律长期以来的承诺。
英文摘要
Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided. We show this is not the case: the confounding factor is hyperparameters. Small models are highly sensitive, but hyperparameter sensitivity fades with scale. This small-scale sensitivity makes scaling laws easy to miss because they only emerge on the fully tuned frontier, and reaching that frontier requires an extensive search far beyond what most ever run. By ablating the basic scaling law recipe, we show well-tuned hyperparameters matter more than any other ingredient. Further, we reveal why those hyperparameters become easier to find: as scale increases, the hyperparameter loss surface becomes lower dimensional. Nevertheless while scaling laws exist in small models, extrapolation hits statistical limitations. A holistic approach is required. Synthesizing our insights with the recent literature, we develop a new methodology for model-centric research and demonstrate it on a question that once took the field years to settle: where to place normalization layers in the transformer architecture. From small-scale experiments, we recover the large scale result: pre-normalization works better as models grow in size. With the right tools and a better understanding, small-scale experiments can deliver on scaling laws' long-awaited promise.