AI 中文总结
研究针对MoE语言模型服务效率问题,提出ExpertPlex系统,通过跨阶段共享专家、解耦注意力模块,结合自适应持久内核等技术,提升模型服务吞吐量,实验显示比现有方法有显著提升。
AI 中文摘要
语言模型通过扩展混合专家(MoE)参数来提升智能,但大量权重和动态计算阻碍了高效服务。现有实例级预填充-解码解耦将各阶段隔离在单独的完整模型副本上,随着MoE权重增加,资源分配变得粗糙。预填充-解码共置可避免重复,但现有绿色上下文解决方案按阶段划分每个GPU并在内核期间固定阶段资源,无法跟踪跨操作的资源变化或路由专家负载的逐层变化。我们提出ExpertPlex,它在解耦轻量级注意力模块的同时跨阶段共享大量MoE专家,消除了超过95%的重复模型权重并复用动态稀疏计算,减少了注意力通信成本。ExpertPlex还使用自适应持久内核、注意力引发的MoE通信和瓦片到集群模型来优化机制以实现最大吞吐量。实验表明,ExpertPlex比实例级预填充-解码解耦的吞吐量提高了2.01倍,比预填充-解码共置提高了1.66倍。
英文摘要
LLMs scale Mixture-of-Experts (MoE) parameters for superior intelligence, but massive weights and dynamic computation impede efficient serving. Existing instance-level prefill-decode disaggregation isolates the phases on separate full-model replicas. As MoE weights grow, each instance may span tens to hundreds of GPUs, making resource allocation increasingly coarse. Configured prefill-to-decode ratios thus often mismatch demand, overprovisioning one phase while overloading the other. Prefill-decode colocation avoids this duplication, but existing Green Context solutions partition each GPU by phase and fix phase resources during a kernel. They cannot track resource changes across operations or layerwise variation in routed expert load, causing head-of-line blocking or idle reserved resources. Partitioning every GPU also leaves each phase with fewer local resources, forces wider parallelism and more communication, and lets prefill and decode traffic interfere on the shared network. We present ExpertPlex, which shares massive MoE experts across phases while disaggregating lightweight attention modules. Expert sharing eliminates over 95% of duplicate model weights and multiplexes dynamically sparse computation, while attention disaggregation reduces attention communication cost. ExpertPlex further uses (1) adaptive persistent kernels to schedule dynamic expert computation at tile granularity for efficient, isolated execution; (2) attention-initiated MoE communication to avoid network interference and enable cross-phase communication-computation overlap; and (3) a tile-to-cluster model to optimize these mechanisms for maximum goodput. Experiments serving MiniMax-M2.7 and GLM-5.1-FP8 show that ExpertPlex improves goodput by up to 2.01$\times$ over instance-level prefill-decode disaggregation and 1.66$\times$ over prefill-decode colocation.