HAPMoE:面向混合专家模型训练的异构感知自动并行规划
HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
- Peking University(北京大学)
- Infinigence AI(无问智能)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对MoE模型在异构集群上的自动并行化难题,提出HAPMoE,通过MoE感知成本模型和六维并行空间搜索,实现吞吐量最高提升3.2倍,并在一分钟内完成规划。
AI中文摘要:
随着模型规模持续扩大,分布式训练已不可避免。自动并行化技术能够以较低成本推导出高效的训练并行策略,同时实现优越的性能。该问题的难度由模型的复杂性和底层计算集群共同决定。与此同时,混合专家(MoE)模型正日益成为主导架构,而加速器硬件的快速演进使得集群异构性成为常态,这对自动并行化提出了重大挑战。然而,现有方法通常要么针对MoE架构,要么针对异构集群,无法推广到两种挑战同时存在的场景。为此,我们提出了HAPMoE,一种面向MoE训练的异构感知自动并行规划器。HAPMoE构建了一个轻量级的MoE感知成本模型,并高效搜索六维并行空间,生成可直接部署在Megatron-LM上的并行方案。实验表明,在异构集群上,HAPMoE相比基线将端到端训练吞吐量最多提升3.2倍。其非均匀流水线划分额外带来最高78%的增益,而其剪枝增强的动态规划算法在1分钟内完成搜索,展示了在复杂硬件环境中的高效率和实用价值。
英文摘要:
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving superior performance. The difficulty of this problem is jointly determined by the complexity of the model and the underlying compute cluster. Meanwhile, mixture-of-experts (MoE) models are increasingly emerging as the dominant architecture and the rapid evolution of accelerator hardware has made cluster heterogeneity commonplace, posing substantial challenges to automatic parallelization. However, existing approaches typically target either MoE architectures or heterogeneous clusters, failing to generalize to scenarios where both challenges coexist. To this end, we present HAPMoE, a heterogeneity-aware automatic parallelism planner for MoE training. HAPMoE builds a lightweight MoE-aware cost model and efficiently searches a six-dimensional parallel space, producing parallel plans directly deployable on Megatron-LM. Experiments show that HAPMoE improves end-to-end training throughput by up to 3.2$\times$ over baselines across heterogeneous clusters. Its non-uniform pipeline partitioning yields an additional up to 78% gains, and its pruning-enhanced dynamic programming algorithm completes the search within 1 minute, demonstrating high efficiency and practical value in complex hardware environments.