发表机构
FAIR at Meta(Meta FAIR实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对神经缩放定律的独立假设缺陷,提出Skaling定律耦合模型容量与数据,可降低MAPE,用更少计算实现准确外推,为模型训练预算分配提供稳健框架。
AI 中文摘要
神经缩放定律是语言模型开发的基础,但标准公式在数据稀缺和过训练极端情况下会系统性地低估和高估损失。这一失败源于其潜在假设:模型规模和训练数据对损失的影响是相互独立的。为解决该问题,我们提出Skaling定律,这是一种通过单个交互指数将模型容量与数据耦合的广义函数形式。这一简单扩展在插值和外推两种情况下,将平均绝对百分比误差(MAPE)降低了1.5至3倍。当与仅限低计算区域的稀疏网格策略结合时,Skaling定律仅需约均匀扫描10倍的计算量即可实现准确的全网格外推。通过支持从小规模实验中进行可靠的性能预测,Skaling定律为下一代模型训练中的计算预算分配提供了更稳健且资源高效的框架。
英文摘要
Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. To address this, we introduce the Skaling law, a generalized functional form that couples model capacity and data through a single interaction exponent. This simple extension reduces the Mean Absolute Percentage Error (MAPE) by 1.5-3x across both interpolation and extrapolation regimes. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law achieves accurate full-grid extrapolation using approximately 10x less compute than uniform sweeps. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.