arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推导OpenEuroLLM模型的缩放定律:学习率、批量大小与损失

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

Niccolò Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, Jörg Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein

arXiv 2608.28308首次发表:更新:

发表机构

Eindhoven University of Technology; Georgia Institute of Technology; University of Freiburg; LAION; Open-Ψ (Open-Sci) Collective; University of Turku; OpenEuroLLM(埃因霍温理工大学; 佐治亚理工学院; 弗赖堡大学; LAION; Open-Ψ(开放科学)集体; 图尔库大学; OpenEuroLLM)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对OpenEuroLLM模型,探究预训练时学习率与批量大小的缩放规律,开发相关模型并评估损失的缩放形式,建立基线与缩放程序并开源预训练运行集合。

AI 中文摘要

我们研究了在以英语为主的语料库上预训练稠密大语言模型时学习率和批量大小的缩放行为。除了缩放联合最优的学习率和批量大小外,我们还探究它们随模型容量和数据规模的边际演化,并开发了一个捕捉这些关系的模型。由于我们采用了 Warmup-Stable-Decay 学习率调度,我们进一步在广泛的超参数设置、模型和数据预算范围内研究了学习率退火带来的增益,以及最优学习率和批量大小是否在稳定阶段和衰减阶段之间迁移。最后,我们描述了损失对模型容量和数据集规模的依赖关系,评估了最近提出的明确建模两者相互作用的缩放形式。我们发现这些方法在实验中对捕捉欠训练和过训练 regime 特别有效。本研究为未来 OpenEuroLLM 模型的开发建立了首个基线和缩放程序,并开源了本研究中使用的完整预训练运行集合。

英文摘要

We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling jointly optimal learning rates and batch sizes, we investigate their marginal evolution with model capacity and data scale and develop a model that captures these relationships. As we employ a Warmup-Stable-Decay learning rate schedule, we further investigate the gains from learning rate annealing over a broad range of hyperparameters settings, models and data budgets, and whether the optimal learning rate and batch size transfer between the stable and decay phases. Finally, we characterize the dependence of loss on model capacity and dataset size, evaluating recently proposed scaling forms that explicitly model their interaction. We find these approaches particularly effective at capturing both undertraining and overtraining regimes across our experiments. This study establishes a first baseline and scaling procedure for the development of future OpenEuroLLM models. We open-source the complete collection of pretraining runs used in this study.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑