专家混合服务
Mixture-of-Experts Serving
浏览论文内容
中文总结 AI 辅助
研究专家混合服务模型中如何动态分配GPU的问题,提出多项式时间\(O(\sqrt{\log k})\)竞争在线算法,给出离线常数因子近似,证明其NP难且排除FPTAS,为MoE服务的资源分配提供了理论算法及复杂度分析。
中文摘要 AI 辅助
专家混合(MoE)模型将每个令牌仅路由到少数专家网络,随着时间推移,服务负载在不同受欢迎程度的专家间分配。服务系统必须动态决定为每个专家分配多少GPU,权衡服务延迟和重新配置分配的成本。我们引入了MoE服务的形式化模型,并对其在线和离线算法进行了有原则的研究。主要成果是一个多项式时间的\(O(\sqrt{\log k})\)竞争在线算法,其中\(k\)是每个专家额外的GPU数量。我们还给出了在线对偶问题的匹配\(\Omega(\sqrt{\log k})\)障碍。在离线设置中,我们给出了常数因子近似,表明MoE服务是NP难的,并假设ETH排除了FPTAS。
英文摘要
Mixture-of-Experts (MoE) models route each token to only a few expert networks, distributing the serving load across experts whose popularity shifts over time. A serving system must therefore dynamically decide how many GPUs to assign to each expert, trading off service latency against the cost of reconfiguring the assignment. We introduce a formal model of MoE Serving and initiate a principled study of online and offline algorithms for it. Our main result is a polynomial-time $O(\sqrt{\log k})$-competitive online algorithm, where $k$ is the number of GPUs beyond one per expert. We complement it with a matching $Ω(\sqrt{\log k})$ barrier for the online dual problem underlying our analysis. In the offline setting, we give a constant-factor approximation, show that MoE Serving is NP-hard, and rule out an FPTAS assuming ETH.