AI 中文总结
本文通过将重复 token 与新数据及同数据一轮训练对比定价,揭示重复成本由额外轮数除以每参数唯一 token 数决定,并给出临界轮数及模型规模与轮数的权衡,为数据受限下的预训练提供经验指南。
AI 中文摘要
随着预训练越来越频繁地重复使用数据,每次运行都面临三个问题:应该进行多少轮(epoch)训练,这个数字应如何随模型规模变化,以及除了轮数之外是否还有其他因素重要。我们通过将重复 token 与两个参照物进行定价来回答这些问题:同一数据上的一轮训练(这给出了其价值),以及在相同计算量下的新数据(这给出了其成本)。与新数据相比,重复的成本由单一变量决定,即额外轮数除以每个参数的唯一 token 数。与同一数据相比,第二轮训练的价值几乎等同于新数据的一轮,而在经过一个临界轮数之后,重复 token 的价值降至新 token 的一半,该临界轮数随每个参数的训练预算增长,但几乎不随模型规模变化。在唯一数据固定的情况下,预测的计算最优运行会同时增加模型规模和轮数,直到损失停止改善,接近临界轮数。同一变量也解释了看似矛盾的规模趋势的方向:当语料库固定时,较大的模型能容忍更少的轮数,从 127M 参数时的约 15 轮降至 2B 参数时的 4 轮,但当唯一数据随模型增长时则不然。仅凭计数并不能决定损失:在相同计数下,连续重放分片会使损失增加高达每字节 0.46 比特,将重复集中在更少的样本上也会增加损失,较低熵的来源在重复时退化更快,而重新分词重复数据仅在重度重复下才有帮助。这些结果为在唯一数据而非计算成为约束条件时的预训练提供了经验指南。
英文摘要
As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters. We answer them by pricing a repeated token against two references: one epoch on the same data, which gives its value, and fresh data at equal compute, which gives its cost. Against fresh data, the cost of repetition follows a single variable, the number of extra epochs divided by the unique tokens per parameter. Against the same data, a second epoch is worth nearly as much as a fresh one, and repeated tokens fall to half the value of fresh ones after a critical epoch count that grows with the training budget per parameter but hardly with model size. With unique data fixed, the predicted compute-optimal run grows model size and epochs together until loss stops improving, near the critical epoch count. The same variable accounts for the direction of size trends that appear to conflict: larger models tolerate fewer epochs when the corpus is fixed, from about 15 at 127M to 4 at 2B parameters, but not when unique data grow with the model. Counts alone do not determine loss: at identical counts, replaying shards consecutively raises loss by up to 0.46~bits per byte, concentrating repeats on fewer samples also raises it, lower-entropy sources degrade faster with repetition, and re-tokenizing repeats helps only under heavy repetition. These results offer an empirical guide to pretraining when unique data, rather than compute, are the binding constraint.