arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为什么自适应批处理有助于大语言模型预训练?来自无界方差的视角

Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance

Arda Fazla, Antesh Upadhyay, Ege C. Kaya, M. Berk Sahin, Abolfazl Hashemi

arXiv 2610.02355首次发表:更新:

发表机构

Elmore Family School of Electrical and Computer Engineering; Purdue University(埃尔莫尔家族电气与计算机工程学院; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM预训练中批大小调度的理论缺失,提出广义BG-$a$噪声模型并推导复杂度界限,设计自适应批调度器,在OLMo2预训练中以更少迭代取得更低验证损失。

AI 中文摘要

在训练期间增加批大小是大语言模型(LLM)预训练中的常见做法,然而其成功背后的理论依据尚未被充分理解。随机优化的分析通常假设随机梯度方差一致有界,但近期证据表明,这一假设在许多实际非凸问题中并不成立。Blum--Gladyshev(BG-$0$)噪声模型通过允许方差随与初始化的距离二次增长来放宽这一假设,这表明批大小调度器可以通过控制训练期间的方差增长来提供帮助。然而,这种增长在实践中可能过于保守。我们实证研究了LLM预训练中的方差增长,观察到具有可调增长指数的广义BG模型能更紧密地描述实际噪声行为。受此观察启发,我们引入了广义BG-$a$噪声模型,该模型在有界方差($a=0$)和BG-$0$噪声($a=2$)之间插值。在$L$-光滑性下,我们推导了一个信息论下界,其具有依赖于增长的预言机复杂度$\Omega(\epsilon^{-(4+a)})$,并通过随着迭代点远离初始化而增加批大小,建立了在$\epsilon$依赖性上的匹配上界。最后,我们提出了一种自适应批调度器,通过在训练期间动态调整批大小来控制方差增长。在C4数据集上预训练参数多达10亿的OLMo2模型时,我们的调度器在匹配的token预算下,比小批大小和大批大小训练都实现了更低的验证损失,同时使用的小批大小训练的迭代次数不到其10%。

英文摘要

Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded stochastic gradient variance, yet recent evidence suggests that this assumption fails in many practical nonconvex problems. The Blum--Gladyshev (BG-$0$) noise model relaxes this assumption by allowing the variance to grow quadratically with the distance from initialization, suggesting that batch size schedulers can help by controlling the variance growth during training. However, this growth can be overly conservative in practice. We empirically investigate variance growth in LLM pretraining and observe that a generalized BG model with a tunable growth exponent provides a tighter description of practical noise behavior. Motivated by this observation, we introduce the generalized BG-$a$ noise model, which interpolates between bounded variance ($a=0$) and BG-$0$ noise ($a=2$). Under $L$-smoothness, we derive an information-theoretic lower bound with growth-dependent oracle complexity $Ω(ε^{-(4+a)})$ and establish a matching upper bound in $ε$-dependence by increasing the batch size as the iterates move away from initialization. Finally, we propose an adaptive batch scheduler that controls variance growth through dynamic batch size adjustments during training. In pretraining OLMo2 models of up to 1B parameters on C4, our scheduler achieves a lower validation loss than both small and large batch training under matched token budgets, while using less than 10\% of the iterations of small batch training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑