ExFold:用于无训练MoE预填充-解码加速的统一专家折叠方法
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
浏览论文内容
中文总结 AI 辅助
ExFold是一种无训练的统一专家折叠框架,作为vLLM插件可联合加速MoE的预填充与解码,最高实现1.41倍TTFT和2.45倍TPOT加速,同时保留约99%的原始平均质量。
中文摘要 AI 辅助
混合专家(Mixture-of-Experts, MoE)模型通过稀疏激活专家,在保持每token计算量受限的同时扩展了强质量所需的容量。然而,低延迟MoE服务正变得越来越具有挑战性,因为它跨越两个推理阶段,具有根本不同的瓶颈:预填充(prefill)阶段以逐token的专家计算为主导,而解码(decode)阶段则受限于批量激活的专家集带来的内存流量。但现有的无训练加速方法仅优化单一资源代理,要么是每个token执行的专家,要么是批量激活的专家,且要么丢弃了被排除专家的贡献,要么仅将其近似隐含处理。本文提出ExFold,一种用于联合加速MoE预填充与解码的无训练统一专家折叠框架。ExFold将预填充与解码均视为一个受预算约束的输出近似问题:仅执行特定阶段的受限专家集,同时通过校准的标量投影器将预算排除专家的贡献投影到保留的专家上。受“许多专家输出在方向上对齐但幅度不同”这一观察的启发,ExFold在未标记数据上校准了成对的标量投影器矩阵,并在推理时用其将被排除专家的贡献折叠到保留的专家中。在此视角下,预填充加速成为token级Top-K折叠,解码加速成为批量级专家池折叠。两个阶段的区别仅在于保留专家的选择方式,而被排除的贡献则通过一种共享的折叠机制恢复。我们将ExFold实现为vLLM中的即插即用插件,配备轻量级专家折叠CUDA内核,实现了最高1.41倍的首token时间(TTFT)加速和2.45倍的每token输出时间(TPOT)加速,同时保留了原始平均质量的约99%。
英文摘要
Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set. However, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts' contribution or leave it only implicitly approximated. In this paper, we propose ExFold, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode. ExFold casts both prefill and decode as one budgeted output-approximation problem: execute only a phase-specific constrained expert set while projecting the contribution of budget-excluded experts onto retained experts using calibrated scalar projectors. Motivated by the observation that many expert outputs are directionally aligned but differ in magnitude, ExFold calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts. Under this view, prefill acceleration becomes token-level Top-K folding, and decode acceleration becomes batch-level expert-pool folding. The two phases differ only in how retained experts are selected, while excluded contributions are recovered by one shared folding mechanism. We implement ExFold as a plug-and-play plugin in vLLM, with a lightweight expert-folding CUDA kernel, delivering up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.