arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoMoE:在消费级GPU上民主化MoE推理

Democratizing MoE inference on commodity GPUs with CoMoE

Ruwen Fan, Yuezhi Zu, Junru Li, Qingda Hu, Xinjun, Yang, Jiwu Shu, Youyou Lu

arXiv 2610.09424首次发表:更新:

发表机构

Tsinghua University; Alibaba Cloud Computing(清华大学; 阿里云计算)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CoMoE通过以主机为中心的路由,在消费级GPU上实现通信高效的MoE推理,减少通信量并消除同步停顿,吞吐量提升1.46倍,成本仅为NVLink方案的23.4%。

AI 中文摘要

部署混合专家(MoE)模型严重依赖专家并行(Expert Parallelism),这会产生大量的GPU间通信。因此,最先进的推理系统需要数据中心GPU中的高带宽、点对点(P2P)互连(如NVLink)来处理海量令牌路由,使得部署成本高昂。消费级GPU以显著更低的成本提供可比较的计算能力,有望使MoE推理民主化,让个人能够使用,并支持隐私保护的本地部署。然而,其带宽受限(仅弱PCIe总线带宽)且由主机中介的互连(不支持P2P)引入了严重的通信瓶颈。我们提出了CoMoE,一种通信高效的MoE推理系统,通过新颖的以主机为中心的路由解决了这一不匹配问题。我们的关键洞察是,独特的通信拓扑提供了将主机提升为主动路由中心的机会,这可以从根本上减少通信量并消除全局同步引起的停顿。具体来说,对于令牌分发,我们引入了主机支持的令牌多播,将共享令牌恰好写入主机一次,消除了出站传输冗余。对于令牌合并,我们提出了一种使用主机暂存缓冲区的细粒度、令牌级聚合机制,取代了僵化的全局同步并减轻了掉队者效应。在RTX 5090 GPU上的评估显示,CoMoE将推理吞吐量提高了最多1.46倍,以仅23.4%的硬件成本接近支持NVLink的A800 GPU的性能。

英文摘要

Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑