arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据受限预训练中跨模型规模的重复次数选择

Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining

Ziyue WANG, T. Kanamori

arXiv 2610.05126首次发表:更新:

发表机构

Institute of Science Tokyo(东京科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究数据受限预训练中重复次数随模型规模的变化,提出候选保留法,利用小模型损失曲线筛选候选值,并给出基于缩放模型的解释。

AI 中文摘要

对于小型语言模型而言最优的重复次数,在更大规模下可能不再是最优的。我们在固定目标数据比例下,将有限的目标语料与通用数据混合进行预训练,研究了这一效应。在基于维基百科的数据和Proof-Pile-2上,实测重复次数的排名随模型规模变化,且一项520M规模的Proof-Pile-2实验证实,将重复次数从十六次减少到八次,在减少训练令牌数的同时改善了损失。我们利用多个较小模型的损失曲线,保留了一小部分有前景的重复次数候选值,以供更大规模评估使用。在PubMed和Caselaw上,在目标模型训练前固定的候选集,在200M和520M两种规模下,均保留了原始评估网格中损失最低的实测次数。这支持将候选保留作为精确点预测的一种实用替代方案。我们还将剪枝回归与一个经验缩放模型联系起来,该模型包含两个相互对立的、依赖重复次数的损失项。对对数模型规模进行一阶展开,得到选择规则所用的线性形式,从而为候选选择过程提供了基于缩放的解释。

英文摘要

The repetition count that works best for a small language model may not remain best at a larger scale. We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction. On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens. We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale. On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M. This supports candidate retention as a practical alternative to exact point prediction. We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms. A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑