arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14306cs.DC

展平长上下文混合专家训练中的每一个内存峰值

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training

Shrey Pandit, Xuan-Phi Nguyen, Yiran Zhao, Shafiq Joty

首次发表
浏览论文内容

中文总结 AI 辅助

针对长上下文MoE训练中四个未受限制的内存峰值,提出四种调度方法(PipelinedLLEP、Ring-DTP、SCO、OffloadStreamAdamW),在保持精确损失和梯度的同时,显著降低峰值内存并提升吞吐量。

中文摘要 AI 辅助

训练一个混合专家(MoE)模型时,在长上下文或大批量场景下,只要任何一个组件的峰值分配超过设备内存,训练就会失败,因此目标是同时处理每一个峰值,而不是平均占用。常见并行方案中有四个峰值未受限制,且各自增长方式不同:专家分派随路由矩阵增长,词汇投影随词元数乘以词汇量增长,梯度检查点边界随深度乘以序列长度增长,优化器状态随参数数量增长。哪一个先耗尽内存取决于模型、上下文长度和设备数量,因此降低最大的峰值只会暴露下一个峰值。我们通过调度将所有四个峰值限制在启动时固定的GPU工作集内:PipelinedLLEP扩展了最不繁忙专家并行,限制每个源对分派块贡献的词元数;Ring-DTP在词汇投影处沿环传递激活或权重分片,并将每个logits块折叠为在线log-sum-exp;选择性检查点卸载(SCO)将每个检查点边界的唯一长寿命张量保留在CPU内存中;OffloadStreamAdamW将优化器卸载的串行CPU Adam更新转变为桶流水线。这四种方法仅改变计算和数据移动的顺序及粒度,因此损失和梯度保持精确。在匹配的组件测试中,它们将MoE分派峰值降低高达59.3%而不损失吞吐量,词汇投影峰值降低86.6%,卸载优化器步骤加速2.05倍。组合应用于120B到667B参数的MoE模型时,它们能在1M上下文长度下训练,是调优后的FSDP2基线可达范围的8到32倍,吞吐量高达其10.4倍。

英文摘要

Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint. Four are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count. Which one runs out first changes with the model, the context length, and the device count, so lowering the largest only exposes the next. We bound all four with schedules whose GPU working set is fixed at launch: PipelinedLLEP extends least-loaded expert parallelism with a cap on the tokens each source contributes to a dispatch chunk, Ring-DTP circulates activations or weight shards around a ring at the vocabulary projection and folds each block of logits into an online log-sum-exp, Selective checkpoint offload (SCO) keeps the one long-lived tensor of each checkpoint boundary in CPU memory, and OffloadStreamAdamW turns the serial CPU Adam update of optimizer offload into a bucket pipeline. All four change only the order and granularity of computation and data movement, so the loss and gradients stay exact. In matched component tests, they cut the MoE dispatch peak by up to $59.3\%$ without losing throughput, the vocabulary projection peak by $86.6\%$, and the offloaded optimizer step by $2.05\times$ faster. Composed on MoE models from 120B to 667B parameters, they train at 1M context length, $8$--$32\times$ the reach of a tuned FSDP2 baseline, and up to $10.4\times$ its throughput.

发表机构

  • Salesforce AI Research(Salesforce 人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑