arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10549cs.PF

Compass:剖析通信与计算算子以实现高效LLM训练

Compass: Dissecting Communication and Computation Operators for Efficient LLM Training

Guangyu Xiang, Lin Zhang, Haoxuan Yu, Xinglin Pan, Shaohuai Shi, Xiaowen Chu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM训练中通信与计算重叠效率低的问题,提出Compass框架,通过双环通信的算子内融合和数学建模的算子间分解优化,实现最高1.42倍端到端加速。

中文摘要 AI 辅助

重叠通信与计算算子是隐藏通信开销、加速GPU集群上大规模语言模型(LLM)训练的常见做法。现有系统通过算子内融合(IntraFusion)或算子间分解(InterDecom)实现这一点,前者将算子打包成单个大内核,后者将张量拆分为多个部分以进行流水线执行。然而,当前的IntraFusion方法未充分利用网络拓扑,导致多GPU系统上带宽使用欠佳,而InterDecom难以确定最佳分解部分数量以实现峰值性能。为解决这些问题,我们提出了Compass,它采用系统优化和全面建模。首先,我们设计了一种新颖的IntraFusion算法,利用双环通信最大化混合NVLink-PCIe系统中的带宽利用率,实现了1.5倍至2.5倍的加速。其次,我们开发了一个分解模型,从数学上推导出InterDecom的最优张量分解度,性能提升最高达1.3倍。最后,我们构建了一个统一的性能框架,能准确确定不同场景下的最佳策略。我们通过288种配置的广泛评估和真实世界应用的端到端实验验证了Compass。结果表明,Compass始终选择最优策略,与Megatron-LM基线相比,端到端加速最高达1.42倍。

英文摘要

Overlapping communication and computation operators is a common practice to hide communication overheads, accelerating large language models (LLMs) training on GPU clusters. Existing systems achieve this through either intra-operator fusion (IntraFusion), which packs operators into a single large kernel, or inter-operator decomposition (InterDecom), which splits a tensor into multiple parts for pipelined execution. However, current IntraFusion methods underutilize network topology, causing suboptimal bandwidth usage on multi-GPU systems, while InterDecom struggles to determine the optimal number of decomposed parts for peak performance. To address these issues, we introduce Compass, which employs systematic optimization and comprehensive modeling. First, we design a novel IntraFusion algorithm leveraging double-ring communications to maximize bandwidth utilization in hybrid NVLink-PCIe systems, achieving 1.5x-2.5x speedups. Second, we develop a decomposition model that mathematically derives the optimal tensor decomposition degree for InterDecom, improving performance by up to 1.3x. Finally, we develop a unified performance framework that accurately determines the best strategy for different scenarios. We validate Compass through extensive evaluation across 288 configurations and end-to-end experiments on real-world applications. The results demonstrate that Compass consistently selects the optimal strategy, achieving up to a 1.42x end-to-end speedup compared to the Megatron-LM baseline.

↑