TrimMoE:面向分布式边缘推理的通信感知自适应深度框架
TrimMoE A communication aware and adaptive depth framework for distributed edge inference
浏览论文内容
中文总结 AI 辅助
TrimMoE是一种面向分布式边缘MoE大模型推理的通信感知自适应深度框架,通过层跳跃、提前退出等技术,在保证任务质量退化不超2%的前提下,将平均延迟降低最多62.8%,减少跨服务器流量与远程执行比例并提升吞吐量。
中文摘要 AI 辅助
在分布式边缘服务器上部署混合专家(Mixture-of-Experts, MoE)大语言模型时,跨服务器的专家传输是主要瓶颈。现有方法主要关注如何更快地访问远程专家,而本文则研究给定层及其后续层是否需要执行。为此,本文提出一种名为TrimMoE的通信感知自适应深度框架,该框架在统一质量预算下,将层跳跃、基于置信度的提前退出与替代执行、服务器-专家选择相结合。具体而言,在离线阶段,TrimMoE冻结主干网络,训练轻量级的每层退出头,校准每层的重要性阈值,并通过感知跳跃/退出的冗余收益分配专家副本。在在线阶段,转换感知的前瞻机制预测令牌的移动,使深度减少针对成本最高的传输;此外,两条反馈规则调整延迟-质量权重和退出阈值。本文证明,替代与跳跃的代理退化始终不超过配置的预算,且仅在校准后的置信门限下才允许提前退出。在包含10台异构服务器的测试平台上,使用Switch-Base-8E、Qwen-MoE-A2.7B和Mixtral-8x7B模型,TrimMoE可将平均延迟降低最多62.8%,减少跨服务器流量和远程执行比例,在负载下保持高吞吐量,同时将任务质量退化控制在2%以内。
英文摘要
Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The existing approaches mainly focus on how to reach a remote expert faster. However, in this paper, we instead consider whether a given layer, and the layers after it, need to be executed at all. To this end, a communication-aware adaptive-depth framework is proposed in this paper, termed TrimMoE, which couples layer skipping and confidence-based early exit with substitute execution and server-expert selection under a unified quality budget. Specifically, in the offline stage, TrimMoE freezes the backbone, trains the lightweight per-layer exit heads, calibrates the per-layer importance thresholds, and allocates the expert replicas by a skip/exit-aware redundancy benefit. In the online stage, a transition-aware look-ahead anticipates the token movement, so that the depth reduction targets the costliest transmissions, and besides, two feedback rules adapt the delay-quality weights and the exit threshold. Moreover, we prove that the substitution-and-skipping proxy degradation never exceeds the configured budget, and that the early exit is admitted only under a calibrated confidence gate. On a heterogeneous 10-server testbed with Switch-Base-8E, Qwen-MoE-A2.7B, and Mixtral-8x7B, TrimMoE reduces the average latency by up to 62.8%, lowers the cross-server traffic and the remote-execution ratio, and sustains high throughput under load, while keeping the task-quality degradation within a 2% bound.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
- Hong Kong University of Science and Technology(香港科技大学)
- University of Science and Technology Beijing(北京科技大学)
机构由 AI 辅助整理,请以论文原文为准。