arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04609cs.DC

CIERA:分片MoE训练中用于无损Allgather的跨迭代指数复用

CIERA: Cross-Iteration Exponent Reuse for Lossless Allgather in Sharded MoE Training

Ali Zafar Sadiq, Haiying Shen, Masahiro Tanaka

首次发表
浏览论文内容

中文总结 AI 辅助

针对分片MoE训练中Allgather通信开销高的问题,提出CIERA无损通信方法,利用跨迭代指数复用减少开销,在OLMoE-1B-7B上实现显著加速且保证参数精确重建。

中文摘要 AI 辅助

在混合专家(Mixture-of-Experts,MoE)模型训练中,分片数据并行将每个专家的参数划分到多个GPU上,需要在每一层执行前进行Allgather操作以重建完整的权重矩阵,该通信过程往往占迭代时间的主要部分。现有研究常采用有损压缩方法减少该开销,但会损失数值精度;而现有的无损方法未利用跨迭代的指数稳定性。本文提出Cross-Iteration Exponent Reuse Allgather(CIERA),一种面向分片MoE训练的无损、系统感知通信方法。我们观察到,经过短暂的预热阶段后,大多数权重的指数值在各迭代间保持不变。基于此,我们将指数本地缓存,仅在指数变化时传输符号和尾数;接收方通过将缓存的指数与接收数据结合,可精确重建原始权重。由于不同层的参数矩阵形状不同,仅在压缩能产生净时间节省时应用压缩。此外,压缩操作与Allgather通信及计算重叠执行。我们的实际实验和大规模轨迹驱动模拟器显示,在16个GPU上的OLMoE-1B-7B模型上,CIERA相比无损基线实现了3.70倍加速,相比有损基线实现了3.68倍加速;在128个GPU上,预计分别达到4.28倍和4.42倍加速,同时在所有评估运行中保持逐位精确的参数重建。

英文摘要

In training Mixture-of-Experts (MoE) models, sharded data parallelism partitions each expert's parameters across GPUs, requiring an Allgather operation to reconstruct the full weight matrix before each layer executes. This communication often dominates iteration time. Prior work often reduces this overhead using lossy compression methods that sacrifice numerical fidelity, while existing lossless methods do not exploit cross-iteration exponent stability. In this paper, we propose Cross-Iteration Exponent Reuse Allgather (CIERA), a lossless, system-aware communication method for sharded MoE training. We observe that after a brief warmup phase, the exponent values of most weights remain unchanged across iterations. Based on this, we cache exponents locally and transmit only the sign and mantissa when exponents are unchanged. The receiver reconstructs the original weights exactly by combining the cached exponents with the received data. Since parameter matrices vary in shape across layers, compression is applied only when it yields net time savings. Moreover, the compression operation is overlapped with both Allgather communication and computation. Our real experiments and large-scale trace-driven simulator show that on OLMoE-1B-7B at 16 GPUs, CIERA achieves a 3.70x speedup over the lossless baseline and 3.68x over the lossy baseline, projected to reach 4.28x and 4.42x respectively at 128 GPUs, while preserving bitwise-exact parameter reconstruction in all evaluated runs.

发表机构

  • University of Virginia(弗吉尼亚大学)
  • Anyscale

机构由 AI 辅助整理,请以论文原文为准。

↑