扩展大语言模型预训练中的领域数据重复策略
Scaling Domain Data Repetition in LLM Pretraining
浏览论文内容
中文总结 AI 辅助
本文研究LLM预训练中领域数据重复的权衡,发现最优重复次数与模型规模弱相关、与领域验证损失强负相关,小代理模型的重复次数可用于估计大模型的情况。
中文摘要 AI 辅助
随着大语言模型(LLM)规模扩大,其训练token预算也需同步增加,以维持合适的每参数token比例(TPP)。然而,高质量领域数据的扩展难度远高于通用网络数据。当模型规模和训练token预算增长时,领域数据在训练混合数据中的占比往往会下降。重复可用的高质量数据是抵消这种稀释效应的有效方法,但过度重复可能导致过拟合。本文在LLM实际扩展场景下研究这一权衡关系,其中训练token预算与模型规模成比例增长。对于固定领域,研究发现:令人惊讶的是,在固定TPP下,最优重复次数随模型规模仅轻微增加;跨不同领域,最优重复次数与该领域的最终验证损失呈强负相关——损失较低的领域通常可从更多重复中获益;相比之下,领域独有数据量与最优重复次数仅呈弱相关。这些发现表明,在相同TPP下,基于较小代理模型调整的重复次数可为更大模型提供实用估计。
英文摘要
As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)). However, high-quality domain data is much harder to scale than general web data. As model size and the training-token budget increase, its fraction in the training mixture tends to decrease. Repeating the available high-quality data provides an effective way to counteract this dilution, but excessive repetition may lead to overfitting. We study this trade-off under practical LLM scaling, where the training-token budget grows proportionally with model size. For a fixed domain, we first find that, surprisingly at a fixed \(\mathrm{TPP}\), the optimal repetition count mildly increases with model size. Across different domains, we find that the optimal repetition count is strongly negatively correlated with the final validation loss of a domain: domains with lower loss can generally benefit from more repetitions. In contrast, the amount of unique domain data is only weakly related to the optimal repetition count. These findings suggest that repetition counts tuned on smaller proxy models with the same \(\mathrm{TPP}\) can provide a practical estimate for larger models.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。