发表机构
NVIDIA Corporation(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出NCCL M2N,一种布局与拓扑感知的集合通信原语,用于分布式张量重分片,通过分层路由和网络传输与本地复制重叠,显著加速DeepSeek-V3训练中的权重同步。
AI 中文摘要
分布式训练和 rollout 生成通常使用不同的张量布局,导致模型权重需要在不同的进程组之间进行重分片。这种 M 对 N 的重新分配无法由标准集合通信直接表达。扁平直接发送会在目标副本之间重复传输流量,而 gather-then-broadcast 则将网络注入集中到单个根节点,并传输目标节点不需要的数据。对于在 256 块 GPU 上运行的 DeepSeek-V3,传统的 all-gather 加 broadcast 操作占报告中的强化学习步骤时间的 29.4%。我们提出了 NCCL M2N,一种面向分布式张量重分片的布局与拓扑感知的集合通信原语。给定源和目标网格及放置,它推导出所需的传输区域和全局通信调度。其分层路由在符合条件的目标 ranks 之间平衡源贡献,在目标 NVLink 域之间转发一份副本,并通过 NVLink 完成本地复制。网络传输与本地复制重叠,避免了副本倍增的源端出口流量。一个聚合的数据移动模型捕获了两个阶段的限制。我们在 NVL72 集群中使用 NDR InfiniBand 在多达 256 块 GB200 GPU 上评估了 NCCL M2N。单个 FFN-MoE 层传输相比扁平直接发送实现了最高 7.9 倍加速(9.8 毫秒对 77.3 毫秒)。在另一个 256-GPU DeepSeek-V3 NeMo-RL 实验中,NCCL M2N 将报告的权重同步时间从 5.78 秒减少到 2.77 秒,相比传统 all-gather 加 broadcast 实现了 2.09 倍加速,并将步骤时间减少了 12.7%。
英文摘要
Distributed training and rollout generation often use different tensor layouts, requiring model weights to be resharded across distinct process groups. This M-to-N redistribution is not directly expressed by standard collectives. Flat direct sends duplicate traffic across destination replicas, while gather-then-broadcast concentrates network injection at one root and transfers data that destinations do not need. For DeepSeek-V3 on 256 GPUs, legacy all-gather plus broadcast accounts for 29.4% of the reported reinforcement-learning step time. We present NCCL M2N, a layout- and topology-aware collective primitive for distributed tensor resharding. Given source and destination meshes and placements, it derives the required transfer regions and a global communication schedule. Its hierarchical route balances source contributions across eligible destination ranks, forwards one copy between destination NVLink domains, and completes local replication over NVLink. Network transfer and local replication overlap, avoiding replica-multiplied source egress. An aggregate data-movement model captures the limits of both stages. We evaluate NCCL M2N on up to 256 GB200 GPUs in an NVL72 cluster with NDR InfiniBand. A single FFN-MoE layer transfer achieves up to 7.9x speedup over flat direct sends (9.8 ms versus 77.3 ms). In a separate 256-GPU DeepSeek-V3 NeMo-RL experiment, NCCL M2N reduces reported weight-sync time from 5.78 s to 2.77 s, a 2.09x speedup over legacy all-gather plus broadcast, and reduces step time by 12.7%.