发表机构
Harvard University; MIT CSAIL(哈佛大学; 麻省理工学院计算机科学与人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对计算增长快于高质量数据可用性的问题,提出计算-数据(CD)缩放定律框架,通过引入令牌有效性函数η拟合数据扩展策略,划分训练操作模式,指出经典计算最优分配在多数实际设置下次优。
AI 中文摘要
经典的计算最优缩放定律假设预训练数据供应无限,但预训练正进入计算增长快于高质量数据可用性的阶段。我们提出了计算-数据(CD)缩放定律,这是一个统一框架,弥合了计算最优缩放(数据随计算自由缩放)和数据最优缩放(语料库固定而计算可无限增长)。CD缩放通过引入令牌有效性函数η扩展了经典缩放定律,η量化派生令牌(如通过多轮重复或释义产生)相对于新令牌的价值。我们使用Dolma-3语料库针对从14M到600M参数的模型大小,对多轮重复和释义这两种数据扩展策略拟合了η。发现令牌有效性并非恒定,它取决于模型大小、参数令牌比和派生数据量,且随着语料库扩展而饱和。η的函数形式意味着随着模型大小或数据可用性增加,用计算替代数据会收益递减。它还将训练分为计算受限、数据受限和模型受限三种操作模式,并表明经典计算最优分配在大多数实际相关设置中是次优的。
英文摘要
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.