arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14205cs.AIcs.LG

FreeBalance:基于残差 workload 预测的预路由在线 MoE 负载均衡

FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction

Pengfei Chen, Yize Wu, Shouxu Kuang, Ke Gao, Ling Li

首次发表
浏览论文内容

中文总结 AI 辅助

FreeBalance 是一种 MoE 模型在线负载均衡框架,通过残差 workload 预测实现专家迁移与预路由计算重叠,降低负载失衡与推理延迟,提升分布式推理效率。

中文摘要 AI 辅助

负载失衡是混合专家(MoE)模型分布式推理中专家并行效率的主要瓶颈。负载最重的 rank 会因路由分布倾斜导致全局执行停滞,直接增加延迟。离线专家放置虽可缓解持续失衡,但实际多任务服务 workload 存在层依赖和批次依赖的路由动态,使得在线负载均衡不可或缺。现有方法依赖 MoE 路由器后收集的路由统计,仅在路由决策可用后才开始专家权重加载或迁移,从而将迁移开销置于推理关键路径上。本研究发现,若能提前准确预测路由分布,在线均衡可在目标路由(如注意力)前的计算中大量重叠。为此,我们提出 FreeBalance,一种无损在线负载均衡框架,通过残差 workload 预测将专家迁移与前序计算阶段重叠。FreeBalance 利用残差网络中隐藏表示的跨层相似性构建轻量级 workload 预测器,支持在路由决策可用前进行主动专家迁移规划,在权重传输与计算密集型预路由阶段间形成大量重叠。此外,成本模型限制交换次数,以在可用窗口内完全隐藏同步开销。跨模型和数据集的实验表明,FreeBalance 将 rank 负载的最大均值比降低 32.8%,端到端预填充延迟降低 13.1%。具体而言,我们的方法每层平均隐藏 5.1 个专家的均衡开销,该开销原本约占关键路径延迟的 8.5%。

英文摘要

Load imbalance poses a major bottleneck to the efficiency of expert parallelism in distributed inference of Mixture-of-Experts (MoE) models. The most heavily loaded rank stalls global execution due to skewed routing distributions, directly increasing latency. While offline expert placement can alleviate persistent imbalance, practical multi-task serving workloads exhibit layer- and batch-dependent routing dynamics, making online load balancing indispensable. Existing approaches rely on routing statistics collected after each MoE router, requiring expert weight load or migration to begin only after routing decisions are available, consequently placing migration overhead on the inference critical path. In this work, we observe that online balancing can instead be largely overlapped with computation before target routing (e.g., attention), if routing distributions can be predicted accurately in advance. Therefore, we propose FreeBalance, a lossless online load-balancing framework that overlaps expert migration with preceding computation stages via residual workload prediction. FreeBalance leverages cross-layer similarities in hidden representations within the residual network to build a lightweight workload predictor. This enables proactive expert migration planning before routing decisions are available, creating substantial overlap between weight transfer and computation-heavy pre-routing stages. Furthermore, a cost model constrains the number of swaps to fully hide the synchronization overhead within the available window. Experiments across models and datasets show that FreeBalance reduces the max-to-mean rank load ratio by 32.8% and end-to-end prefill latency by 13.1%. Specifically, our method hides balancing overhead of an average of 5.1 experts per layer, which would otherwise account for about 8.5% of the critical-path latency.

补充信息

↑