TopoCompress:面向分布式边缘MoE推理的拓扑感知令牌压缩算法
TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference
- University of Science and Technology Beijing(北京科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出TopoCompress框架,联合优化令牌压缩、专家部署与路由,通过双时间尺度交替优化降低边缘MoE推理的跨服务器通信与资源消耗,同时保持推理质量。
AI中文摘要:
混合专家(MoE)模型通过按令牌稀疏激活专家,以适度的开销提升模型容量。然而,由于专家分布在异构服务器上,将MoE部署在资源受限的边缘服务器上会产生大量的跨服务器通信。现有的放置方法优化原始令牌流量,而传统压缩方法考虑语义却忽略了依赖拓扑的路由成本。因此,独立优化会导致通信和资源利用效率低下。本文提出TopoCompress,一种部署与拓扑感知的令牌压缩框架,用于通信高效的分布式边缘MoE推理。它联合优化令牌压缩、专家部署/复制、GPU-CPU驻留以及协作路由,以平衡跨服务器传输、推理质量和资源使用。为解决令牌级压缩与周期级部署之间的耦合问题,TopoCompress采用双时间尺度交替优化。在在线快速循环中,它识别并压缩低重要性、高路由成本的令牌,并联合路由存活的专家激活。在离线慢速循环中,它根据在线推理期间累积的压缩后流量更新专家放置、复制和GPU-CPU驻留。我们论证了其可行性、最优性、收敛性和计算复杂度。仿真表明,TopoCompress有效减少跨服务器流量和部署资源消耗,同时保持可控的推理质量,从而在带宽和资源受限的边缘基础设施上实现高效的分布式MoE推理。
英文摘要:
Mixture-of-experts (MoE) models improve capacity with moderate overhead by sparsely activating experts per token. However, deploying MoE across resource-constrained edge servers incurs substantial cross-server communication as experts are distributed across heterogeneous servers. Existing placement methods optimize for raw token traffic, while conventional compression considers semantics but ignores topology-dependent routing costs. Consequently, independent optimization leads to inefficient communication and resource utilization. This paper proposes TopoCompress, a deployment- and topology-aware token compression framework for communication-efficient distributed edge MoE inference. It jointly optimizes token compression, expert deployment/replication, GPU-CPU residency, and collaborative routing to balance cross-server transmission, quality, and resource use. To address the coupling between token-level compression and epoch-level deployment, TopoCompress employs a two-timescale alternating optimization. In the online fast loop, it identifies and compresses low-importance, high-routing-cost tokens and jointly routes surviving expert activations. In the offline slow loop, it updates expert placement, replication, and GPU-CPU residency according to post-compression traffic accumulated during online inference. We establish the feasibility, optimality, convergence, and computational complexity. Simulations demonstrate that TopoCompress effectively reduces cross-server traffic and deployment resource consumption while maintaining controllable inference quality, enabling efficient distributed MoE inference over bandwidth- and resource-constrained edge infrastructures.