arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PCIe连接消费级GPU上的高效专家并行通信

Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs

Jaehwan Lee, Sangmin Lee, Chaewon Kim, Junsik Shin, Jaejin Lee

arXiv 2609.40093首次发表:更新:

发表机构

Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对PCIe连接消费级GPU的MoE推理,提出ThunderEP通信设计,消除中继跳、利用DMA引擎并降低轮询开销,在RTX 4090/5090系统上实现最高1.66倍端到端加速。

AI 中文摘要

专家并行(EP)通过将大型混合专家(MoE)模型的专家分布在多个GPU上实现推理,但在每个MoE层都需要大量GPU间通信。随着当代MoE模型每个token激活更多专家,这种通信在推理时间中占比越来越大。在基于PCIe的消费级GPU系统中,成本尤为突出,因为所有GPU间传输都经过CPU内存。然而,现有的MoE专用EP通信库假设直接GPU到GPU访问可用,很大程度上忽视了消费级GPU。因此,大多数LLM框架转而依赖NCCL,其CPU暂存通信导致冗余PCIe传输,并与专家计算竞争GPU资源,限制了重叠。我们提出ThunderEP,一种针对此类系统的新型通信设计,消除了传统环形算法的中继跳,通过DMA引擎移动数据以避免计算资源争用,并通过减少CPU内存中完成标志的轮询开销来最小化同步延迟。我们将该设计集成到vLLM中,并在三个广泛使用的MoE模型上评估。在配备RTX 4090和RTX 5090 GPU的两个PCIe系统上的实验表明,ThunderEP在分发和合并方面相比NCCL分别实现平均2.00倍和1.53倍的加速,端到端相比最先进的MoE推理框架最高可达1.66倍加速。

英文摘要

Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of inference time. The cost becomes particularly pronounced on PCIe-based consumer GPU systems, where all inter-GPU transfers traverse CPU memory. However, existing MoE-specialized EP communication libraries assume that direct GPU-to-GPU access is available, largely overlooking consumer GPUs. Therefore, most LLM frameworks instead rely on NCCL, whose CPU-staged communication incurs redundant PCIe transfers and competes with expert computation for GPU resources, limiting their overlap. We present ThunderEP, a novel communication design for such systems that removes the relay hops of traditional ring algorithm, moves data through DMA engines to avoid compute resource contention, and minimizes synchronization latency by reducing the polling overhead of completion flags in CPU memory. We integrate the proposed design into vLLM and evaluate it on three widely used MoE models. Experiments on two PCIe systems equipped with RTX 4090 and RTX 5090 GPUs show that ThunderEP achieves average speedups of 2.00$\times$ and 1.53$\times$ over NCCL for dispatch and combine, respectively, and up to 1.66$\times$ end-to-end speedup over state-of-the-art MoE inference frameworks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑