AI 中文总结
EasyBalance是一种无需修改专家-设备映射的跨层负载均衡策略,通过联合执行跨层MoE工作负载缓解不平衡,可减少GPU空闲超40%,加速分布式MoE推理。
AI 中文摘要
负载均衡已成为混合专家(Mixture-of-Experts,MoE)模型的专家并行分布式推理中的关键问题。由于专家间的路由分布通常是倾斜的,承载轻负载专家的设备必须空闲等待负载最重的专家完成计算,从而导致效率低下。现有的负载均衡方法主要依赖于每层内的专家复制或迁移,这会引入额外开销并限制其灵活性和可扩展性。为解决该问题,我们提出EasyBalance,一种跨层负载均衡策略,无需修改专家-设备映射,即可实现即时适应性且几乎不产生额外开销。我们的核心见解是:(1)其他层的专家可视为当前层的天然冗余资源;(2)跨层MoE工作负载可联合执行以缓解各自的不平衡。基于这些观察,EasyBalance在每个MoE步骤中贪婪地调度一部分跨层工作负载运行,并将剩余工作负载推迟到后续平衡机会处理,有效利用跨层不平衡缓解。在不同模型、任务和配置上进行的大量实验表明,EasyBalance可持续加速分布式MoE推理,减少GPU空闲时间,幅度大多超过40%。代码可在该https URL获取。
英文摘要
Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. As routing distributions are typically skewed across experts, devices hosting lighter-loaded experts must idle to wait for the heaviest during expert computing, leading to inefficiency. Existing load-balancing approaches primarily rely on expert replication or migration within each layer, which introduce additional overhead and limit their flexibility and scalability. To address this problem, we propose EasyBalance, a cross-layer load balancing strategy that requires no modifications to the expert-device mapping, enabling instant adaptability and incurring essentially no additional overhead. Our key insights are that (1) experts of other layers can be viewed as naturally redundant for the current layer, and (2) cross-layer MoE workloads can be jointly executed to mitigate their individual imbalance. Based on these observations, EasyBalance greedily schedules a subset of cross-layer workloads to run at each MoE step and defers the remaining workloads for future balancing opportunities, effectively leveraging cross-layer imbalance mitigation. Extensive experiments across models, tasks, and configurations demonstrate that EasyBalance consistently accelerates distributed MoE inference, reducing GPU idling by mostly over 40%. Code is available at https://github.com/yize-wu/EasyInfra.
Comments14 pages, 20 figures, ICML 2026