发表机构
Stanford University; Cursor Research(斯坦福大学; Cursor研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对scale-up架构中MoE训练性能不佳的问题,提出Mixture-of-Kittens系统,通过通信方式选择、重叠重构和消除CPU-GPU同步,实现高达2.37倍吞吐量提升。
AI 中文摘要
AI加速器系统正迅速整合为scale-up架构,其中数十至数千个GPU通过高带宽、单跳网络进行通信。我们发现,现有的混合专家(MoE)训练系统,为传统scale-out网络优化,在此设置下迁移效果不佳,通常比用PyTorch和NCCL构建的朴素基线运行得更慢。随着行业路线图指向更大的scale-up领域,理解该硬件体制的性能权衡日益重要。我们提出Mixture-of-Kittens(MoK),一个为Nvidia NVL72设计的MoE训练系统。MoK基于三个见解以解锁scale-up领域的性能:(1)为每个算子选择基于推送或拉取的通信方式,(2)重构计算-通信重叠,(3)完全消除CPU-GPU同步。MoK将这些见解提炼为一个确定性的训练巨型内核,融合了令牌分发、共享和路由专家前馈网络(FFN)以及令牌合并。在来自四个广泛使用的开源模型的MoE层形状中,MoK提供了高达$2.37\ imes$的最强公开可用基线的吞吐量。在跨越多个GB300 NVL72机架的512个GPU的生产运行中,MoK将端到端训练吞吐量提高了$1.41\ imes$。
英文摘要
AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single-hop fabrics. We find that existing Mixture-of-Experts (MoE) training systems, optimized for conventional scale-out networks, transfer poorly to this setting, often running slower than a naive baseline built with PyTorch and NCCL. With industry roadmaps pointing toward even larger scale-up domains, understanding the performance tradeoffs of this hardware regime is increasingly important. We present Mixture-of-Kittens (MoK), an MoE training system designed for Nvidia NVL72. MoK builds on three insights that unlock performance on scale-up domains: (1) choosing push- or pull-based communication per operator, (2) restructuring the computation-communication overlap, and (3) fully eliminating CPU-GPU synchronization. MoK distills these insights into a single deterministic training megakernel that fuses token dispatch, shared and routed expert FFNs, and token combine. Across MoE layer shapes from four widely used open-weight models, MoK delivers up to $2.37\times$ the throughput of the strongest publicly available baseline. In a production run on 512 GPUs spanning multiple GB300 NVL72 racks, MoK improves end-to-end training throughput by $1.41\times$.