arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更大的批次何时有助于扩展大语言模型的强化学习?

When Do Larger Batches Help Scale LLM Reinforcement Learning?

Ziniu Li, Jinbo Wang, Guanhua Huang, Feiyuan Zhang, Pengbo Li, Alex Chen

arXiv 2608.29296首次发表:更新:

发表机构

Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型强化学习,通过分离算法与系统效应,提出更大批次缩短目标时间的决策规则,结合GRPO、PPO实验验证了学习率调整与批次优化可缩短目标时间。

AI 中文摘要

更大的批次会降低每次更新的随机梯度方差,因此通常被认为可以加速训练。然而,这种统计优势是否能转化为达到目标性能的 wall-clock time(墙钟时间)缩短仍不明确,因为每次更新会消耗更多样本,且执行时间可能更长。我们在大语言模型的强化学习中研究这种权衡,通过沿算法和系统效应的自然轴进行比较,将两者分离。在算法层面,我们在累计样本数相等的情况下调整与批次相关的超参数,在一定的批次大小范围内,该过程会产生近似批次大小不变的配置族,其成员遵循相似的以样本为索引的学习轨迹。在系统层面,我们利用 rollout(轨迹生成)与训练之间的计算不对称性:自回归生成在低并发度下通常受内存带宽限制,而训练工作量近似与处理的 token 数成正比。结合这两种观点可得到直接决策规则:更大的批次配置仅当它的吞吐量增益超过其达到目标所需的样本数惩罚时,才会缩短达到目标的时间。使用 GRPO 和 PPO 的实验支持这种分解的两方面:在算法层面,Adam 的平方根学习率缩放在一定批次大小范围内产生近似批次大小不变的学习曲线;在系统层面,更大的批次在固定硬件上将生成吞吐量提高了多达 2.29 倍。在 GRPO 中,将更高的吞吐量与学习率调整相结合可将达到目标的时间缩短多达 29%,而不调整学习率仅增大批次则尽管吞吐量更高,速度反而更慢。

英文摘要

Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. Yet whether this statistical benefit translates into lower wall-clock time-to-target remains unclear, because each update consumes more samples and may take longer to execute. We study this tradeoff in reinforcement learning for large language models. We separate its algorithmic and systems effects by comparing learning and execution along their natural axes. At the algorithmic level, we compare configurations at equal cumulative sample counts while retuning batch-dependent hyperparameters. Over a bounded range of batch sizes, this procedure yields an approximately batch-size-invariant family whose members follow similar sample-indexed learning trajectories. At the systems level, we exploit the computational asymmetry between rollout generation and training: autoregressive generation is often memory-bandwidth-bound at low concurrency, whereas training work scales approximately with the number of processed tokens. Combining these two views yields a direct decision rule: a larger-batch configuration reduces time-to-target only when its throughput gain exceeds its samples-to-target penalty. Experiments with GRPO and PPO support both sides of this decomposition. At the algorithmic level, square-root learning-rate scaling with Adam produces approximately batch-size-invariant learning curves over a bounded range of batch sizes. At the systems level, larger batches improve generation throughput by up to 2.29x on fixed hardware. In GRPO, combining higher throughput with learning-rate retuning reduces time-to-target by up to 29%, whereas increasing the batch without retuning is slower despite its higher throughput.

Comments16 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑