发表机构
Meta; National University of Singapore; Atomic Machines(Meta; 新加坡国立大学; 原子机器公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出幂律熵搜索(PLES),一种基于多保真度贝叶斯优化的获取函数,可高效估计LLM的最优超参数缩放定律,其计算预算仅为传统网格搜索的十分之一。
AI 中文摘要
最优超参数缩放定律描述了大语言模型(LLM)训练的最佳超参数如何随模型和数据规模变化,使从业者无需昂贵的大规模调优即可预测生产规模下的最优配置。然而,传统估计这些缩放定律需要对数千次训练运行进行穷尽网格搜索,消耗巨大计算资源。我们提出幂律熵搜索(Power-Law Entropy Search, PLES),一种基于多保真度贝叶斯优化的计算成本感知获取函数,通过自适应实验高效估计最优超参数缩放定律。PLES的关键创新在于,它搜索能降低缩放定律估计整体不确定性的候选,而非优化单一目标函数。每次迭代中,PLES选择单位计算成本下最大程度降低缩放定律估计不确定性的候选配置,自然倾向于提供信息的小规模实验。我们在合成基准、拟合真实LLM训练数据的代理模型以及实际LLM预训练运行中评估PLES,在所有设置下,PLES收敛到准确的最优超参数缩放定律所用计算预算不到传统网格搜索及其他基线的十分之一。
英文摘要
Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expensive large-scale tuning. However, estimating these scaling laws conventionally requires exhaustive grid searches over thousands of training runs, consuming enormous computational resources. We introduce Power-Law Entropy Search (PLES), a computational cost-aware acquisition function built on multi-fidelity Bayesian optimization that efficiently estimates optimal hyperparameter scaling laws through adaptive experimentation. A key innovation in PLES is that it searches for candidates that reduce the overall uncertainty of a scaling law estimate, instead of optimizing a single objective function. At each iteration, PLES selects the candidate configuration that maximally reduces the uncertainty of the scaling law estimates per unit computational cost, naturally favoring informative small-scale experiments. We evaluate PLES on synthetic benchmarks, surrogate models fitted to real LLM training data, and actual LLM pre-training runs. Across all settings, PLES converges to accurate optimal hyperparameter scaling laws using less than one-tenth of the computational budget required by conventional grid search and other baselines.