arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HetRoute:面向分布式边缘MoE推理的异构感知成本协作路由框架

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference

Xin Yuan, Ning Li, Wenchao Xu, Song Guo, Haijun Zhang

arXiv 2608.00577首次发表:更新:

发表机构

Hong Kong University of Science and Technology; University of Science and Technology Beijing(香港科技大学; 北京科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对分布式边缘MoE推理的路由挑战,提出HetRoute框架,通过统一成本模型优化部署与在线路由,在降低延迟、流量的同时提升吞吐量,且质量损失可控。

AI 中文摘要

混合专家(Mixture-of-Experts,MoE)模型已成为大规模AI服务的主流架构,但将其部署在地理分布式的异构边缘服务器上仍面临挑战。当一个token的Top-k激活专家分布在多个服务器时,最优路由需同时考虑跨服务器链路带宽、异构GPU计算能力、GPU-CPU专家加载延迟、瞬时队列积压以及副本级量化质量损失。现有的分布式推理和MoE服务方法分别处理这些因素,未提供统一的在线多服务器协作路由框架。本文提出HetRoute,一种面向分布式边缘MoE推理的异构感知成本协作路由框架。HetRoute引入统一的逐分配成本模型,明确捕捉四类成本分量:跨服务器传输、GPU-CPU卸载、带队列的GPU计算以及量化诱导的质量惩罚。在该模型指导下,离线阶段通过路由-成本耦合部署算法确定专家服务器放置、GPU-CPU驻留及副本精度;在线阶段通过精确枚举或束搜索最小化瓶颈层成本,将Top-k激活专家集作为整体进行路由。理论分析确立了回退可行性、参与服务器数量的边界、小候选域的每层最优性以及在线计算复杂度。在由10台异构服务器组成的边缘测试平台上,针对三个MoE模型的轨迹驱动评估显示,HetRoute与代表性基线相比,平均推理延迟最多降低59.0%,P99延迟最多降低58.0%,跨服务器流量最多减少72.1%,吞吐量提升2.13倍,同时将质量下降控制在配置预算内。

英文摘要

Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challenging. When the Top-k activated experts of a token are spread across multiple servers, the optimal routing depends jointly on cross-server link bandwidth, heterogeneous GPU computing capability, GPU-CPU expert loading delay, instantaneous queueing backlog, and replica-level quantization quality loss. Existing distributed inference and MoE serving methods address these factors separately and do not provide a unified framework for online multi-server collaborative routing. In this paper, we propose HetRoute, a heterogeneous-cost-aware collaborative routing framework for distributed edge MoE inference. HetRoute introduces a unified per-assignment cost model that explicitly captures four cost components: cross-server transmission, GPU-CPU offloading, GPU computation with queueing, and quantization-induced quality penalty. Guided by this model, the offline stage determines expert server placement, GPU-CPU residency, and replica precision through a routing-cost-coupled deployment algorithm, while the online stage routes the Top-k activated expert set as a whole by minimizing the bottleneck layer cost via exact enumeration or beam search. Theoretical analysis establishes fallback feasibility, a bound on the number of participating servers, per-layer optimality for small candidate domains, and online computational complexity. Trace-driven evaluation on three MoE models over a heterogeneous 10-server edge testbed shows that HetRoute reduces average inference latency by up to 59.0% and P99 latency by up to 58.0%, cuts cross-server traffic by up to 72.1%, and achieves 2.13x throughput improvement compared with representative baselines, while keeping quality degradation within the configured budget.

Comments15 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑