AI 中文总结
研究在直接连接拓扑上实现高效MoE的方法,核心方法是MoX构建令牌感知多播树并利用静态链路权重平衡流量,主要贡献是加速MoE块,在随机拓扑中性能近乎理想,在特定模型上降低调度瓶颈链路负载。
AI 中文摘要
光交换网络适合密集ML模型的常规通信,但混合专家(MoE)会引入稀疏的、运行时相关的流量。我们表明,高效的离线优化路由能够在直接连接拓扑上实现高效的MoE训练和推理,而无需MoE流量矩阵或动态拓扑重新配置。MoX构建令牌感知多播树以减少带宽开销,然后通过解决受限多播树打包问题,使用静态的、预先计算的链路权重来平衡流量。通过使用来自大型MoE模型的记录流量、令牌级跟踪和ASTRA-sim,我们发现MoX在整个MoE块(调度、专家计算和合并)上比最小跳路由加速高达1.8倍。此外,它在随机扩展器拓扑中实现了近乎理想的分组交换网络性能。在谷歌Boardfly拓扑的1024个TPU模型上,MoX将调度瓶颈链路负载降低了高达47%。这些结果表明,通过优化的负载无关路由可以在静态直接连接结构上实现高性能的MoE,而无需按需驱动的重新配置。
英文摘要
Optically switched networks suit the regular communication of dense ML models, but MoE introduces sparse, runtime-dependent traffic. We show that efficient offline-optimized routing enables efficient MoE training and inference on direct-connect topologies without the need for MoE traffic matrix or dynamic topology reconfiguration. MoX constructs token-aware multicast trees to reduce bandwidth tax, then uses static, precomputed link weights to balance traffic by solving a restricted multicast tree-packing problem. Using recorded traffic from large MoE models, token-level traces, and ASTRA-sim, we find that MoX accelerates the full MoE block -- dispatch, expert computation, and combine -- by up to 1.8x over min-hop routing. Moreover, it attains nearly ideal packet-switched network performance in random expander topologies. On a 1,024-TPU model of Google's Boardfly topology, MoX reduces the dispatch bottleneck link load by up to 47%. These results show that high-performance MoE on static direct-connect fabrics can be achieved via optimized load-oblivious routing without demand-driven reconfiguration.
Comments8 pages, 6 figures