AI 中文总结
针对MoE服务中AFD架构下专家负载不均问题,提出AFORE系统,利用AFD的预知与流水线重叠特性进行及时专家重配置,显著提升吞吐并降低延迟。
AI 中文摘要
混合专家(MoE)模型的高效服务因专家参数庞大、专家激活依赖于输入以及工作负载动态变化而面临挑战。专家并行将专家计算分布到多个GPU上,而注意力-FFN分离(AFD)则将注意力与前馈计算分离到独立的工作池中。然而,我们观察到,朴素的AFD实现可能使专家负载不均衡问题更加严重:一旦FFN计算成为独立的流水线阶段,过载的专家会直接拖慢FFN阶段,降低端到端服务性能。为解决此问题,我们提出了AFORE,一个面向基于AFD的MoE服务的及时专家重配置系统。AFORE利用AFD的两个架构特性。第一,AFD在即将到来的微批次到达FFN执行之前就暴露其专家-令牌分布,从而能够基于近期未来需求而非过时的历史配置文件做出放置决策。第二,AFD创建了一个流水线窗口,在该窗口中,针对目标微批次的专家迁移可以与先前在途微批次的计算重叠进行。AFORE将专家重配置形式化为一个微批次感知的调度问题,并使用一个迁移感知的调度器来决定何时迁移以及迁移哪些专家。AFORE进一步实现了轻量级的需求预取和基于NVLink的GPU到GPU专家迁移,以减少重配置开销。在110B参数的MoE模型上,针对四个动态工作负载的评估显示,与最强的竞争基线相比,AFORE将输出吞吐量提高了10.1%-17.6%,并将P95令牌间延迟降低了7.1%-9.5%。与静态放置相比,AFORE平均将吞吐量提高了29.8%,并将P95令牌间延迟平均降低了18.2%。迁移分析进一步表明,AFD流水线重叠可以完全隐藏专家迁移延迟。
英文摘要
Efficient serving of Mixture-of-Experts (MoE) models is challenging due to large expert parameters, input-dependent expert activation, and dynamic workloads. Expert parallelism distributes expert computation across GPUs, while attention-FFN disaggregation (AFD) separates attention and feed-forward computation into independent worker pools. However, we observe that a naive AFD implementation could make expert load imbalance more harmful: once FFN computation becomes an independent pipeline stage, overloaded experts directly slow the FFN stage and degrade end-to-end serving performance. To solve this problem, we present AFORE, a timely expert reconfiguration system for AFD-based MoE serving. AFORE exploits two architectural properties of AFD. First, AFD exposes the expert-token distribution of upcoming microbatches before they reach FFN execution, enabling placement decisions based on near-future demand instead of stale historical profiles. Second, AFD creates a pipeline window in which expert migration for a target microbatch can be overlapped with the computation of preceding in-flight microbatches. AFORE formulates expert reconfiguration as a microbatch-aware scheduling problem and uses a migration-aware scheduler to decide when and which experts to migrate. AFORE further implements lightweight demand prefetching and NVLink-based GPU-GPU expert migration to reduce reconfiguration overhead. Evaluation on a 110B-parameter MoE model across four dynamic workloads shows that AFORE improves output throughput by 10.1-17.6% and reduces P95 inter-token latency by 7.1-9.5% compared with the strongest competing baseline. Compared with static placement, AFORE improves throughput by 29.8% on average and reduces P95 inter-token latency by 18.2% on average. Migration profiling further shows that AFD pipeline overlap can fully hide expert-migration latency.