arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弥合计算最优和数据最优预训练之间的差距

Bridging Compute- and Data-Optimal Pretraining

Tian Qin, Kimia Hamidieh, David Alvarez-Melis

arXiv 2607.25271首次发表:更新:

发表机构

Harvard University; MIT CSAIL(哈佛大学; 麻省理工学院计算机科学与人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对计算增长快于高质量数据可用性的问题,提出计算-数据(CD)缩放定律框架,通过引入令牌有效性函数η拟合数据扩展策略,划分训练操作模式,指出经典计算最优分配在多数实际设置下次优。

AI 中文摘要

经典的计算最优缩放定律假设预训练数据供应无限,但预训练正进入计算增长快于高质量数据可用性的阶段。我们提出了计算-数据(CD)缩放定律,这是一个统一框架,弥合了计算最优缩放(数据随计算自由缩放)和数据最优缩放(语料库固定而计算可无限增长)。CD缩放通过引入令牌有效性函数η扩展了经典缩放定律,η量化派生令牌(如通过多轮重复或释义产生)相对于新令牌的价值。我们使用Dolma-3语料库针对从14M到600M参数的模型大小,对多轮重复和释义这两种数据扩展策略拟合了η。发现令牌有效性并非恒定,它取决于模型大小、参数令牌比和派生数据量,且随着语料库扩展而饱和。η的函数形式意味着随着模型大小或数据可用性增加,用计算替代数据会收益递减。它还将训练分为计算受限、数据受限和模型受限三种操作模式,并表明经典计算最优分配在大多数实际相关设置中是次优的。

英文摘要

Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑