发表机构
ELLIS Institute Tübingen; OpenEuroLLM; Eindhoven University of Technology; University of Freiburg; Zuse School ELIZA; Max Planck Institute for Intelligent Systems; Tübingen AI Center(埃利斯研究所蒂宾根分所; OpenEuroLLM; 埃因霍温理工大学; 弗赖堡大学; 楚塞ELIZA学校; 马克斯·普朗克智能系统研究所; 蒂宾根人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对推导大型基础模型缩放定律计算成本高的问题,提出将数据收集转化为贝叶斯优化的框架,结合逐步扩大计算预算与代理幻想评估,可在节省10至100倍计算成本的同时实现准确的缩放定律拟合。
AI 中文摘要
缩放定律指导大型基础模型训练的设计选择,但推导它们需要对超参数、令牌预算和参数数量进行详尽的网格搜索,计算成本高昂。然而,拟合缩放定律仅需要不同计算规模下的最佳损失前沿,可丢弃大部分训练过的配置。我们提出一种高效构建缩放定律的框架,将数据收集表述为贝叶斯优化问题,并引入在受限计算预算下比较缩放定律拟合方法的指标。我们发现,在采集过程中逐步扩大计算预算(与实践中按计算顺序评估配置的方式一致)可显著提升恢复效率;用代理幻想评估增强观测到的配置则可恢复更广泛的实验网格,无需训练每个配置即可实现准确的缩放定律拟合。这些方法结合后,能以10至100倍的计算成本节省,达到与完整密集网格训练得到的缩放定律相近的拟合效果。
英文摘要
Scaling laws guide the design choices for training large foundation models, but deriving them involves training an exhaustive grid over hyperparameters, token budgets, and parameter counts, which is computationally expensive. Fitting a scaling law, however, only requires the best-loss frontier across compute scales, discarding most of the trained configurations. We propose a framework for efficient scaling law construction that formulates data collection as a Bayesian optimization problem, and introduce metrics for comparing scaling law fitting methods under constrained compute budgets. We find that progressively expanding the compute budget during acquisition, mirroring the compute-ordered evaluation of configurations in practice, substantially improves recovery efficiency. Augmenting the observed configurations with surrogate-fantasized evaluations then recovers the broader experimental grid, allowing accurate scaling law fitting without training every configuration. Together, these can closely match scaling law fits over a full dense grid at computational savings of up to $10\text{--}100\times$.
Comments5 pages, 2 figures, workshop