发表机构
Hong Kong University of Science and Technology; Universidad Politécnica de Madrid; Guangdong University of Technology(香港科技大学; 马德里理工大学; 广东工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在资源受限边缘基础设施部署MoE推理的挑战,提出OrderMoE框架,通过构建专家相似度模型、设计分组部署策略和选择算法,平衡推理延迟等指标,实验表明其能显著降低相关指标,仅带来小的推理质量下降。
AI 中文摘要
尽管混合专家(MoE)模型已越来越多地被用于以适度的计算成本扩展大型语言模型,但在资源受限和带宽有限的边缘基础设施上部署MoE推理仍然具有挑战性。现有分布式MoE服务方法主要依赖精确的专家放置、缓存、复制或通信调度,而忽略了专家之间的功能相似性,这为减少跨服务器令牌传输提供了机会。因此,本文介绍了一种相似度感知专家分配和分布式部署框架OrderMoE,旨在加速边缘MoE推理,同时平衡推理延迟、通信开销、服务器工作负载和推理质量。OrderMoE首先基于路由器诱导的逻辑表示构建专家相似度模型,并将每个MoE层中的专家划分为多个相似度组。然后,它开发了一种相似度感知专家分组和部署策略,以提高边缘服务器之间的局部相似度覆盖。由于减少远程专家调用和保持精确推理质量是相互冲突的目标,OrderMoE进一步设计了一种质量感知和轨迹感知的运行时服务器-专家选择算法,以决定一个令牌是否应该调用其远程目标专家或使用可行的本地替代专家。在真实分布式边缘测试平台上的实验结果表明,OrderMoE显著降低了平均延迟、尾部延迟、跨服务器流量和远程专家调用率,同时仅引入了小的且可控的推理质量下降。
英文摘要
Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to deploy MoE inference over resource-constrained and bandwidth-limited edge infrastructures. Existing distributed MoE serving methods mainly rely on exact expert placement, caching, replication, or communication scheduling, while overlooking the functional similarity among experts, which provides an opportunity to reduce cross-server token transmission. Therefore, this paper introduces a similarity-aware expert allocation and distributed deployment framework, dubbed OrderMoE, which aims to accelerate edge MoE inference while balancing inference latency, communication overhead, server workload, and inference quality. OrderMoE first constructs an expert similarity model based on router-induced logits representations and partitions experts in each MoE layer into multiple similarity groups. Then, it develops a similarity-aware expert grouping and deployment strategy to improve local similarity coverage across edge servers. Since reducing remote expert invocation and preserving exact inference quality are conflicting objectives, OrderMoE further designs a quality-aware and trajectory-aware runtime server-expert selection algorithm to decide whether a token should invoke its remote target expert or use a feasible local substitute expert. Experimental results on a real distributed edge testbed show that OrderMoE significantly reduces average latency, tail latency, cross-server traffic, and remote expert invocation ratio, while introducing only small and controllable inference quality degradation.
Comments17 pages, 12 figures