arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34334cs.NI

HOCCL:将集体通信从GPU核心卸载以加速分布式训练

HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training

Yao Fei, Gongming Zhao, Hongli Xu, Jin Fang, Jiacheng Zhu, Shuo Xu, Kun Huang, Zhuolong Yu

首次发表
浏览论文内容

中文总结 AI 辅助

HOCCL通过DMA引擎驱动通信,实现零SM集体通信框架,消除SM占用,平均性能达最先进水平3%以内,并将训练吞吐量提升最多5%。

中文摘要 AI 辅助

大语言模型训练涉及在GPU流式多处理器(SM)上进行大规模计算,SM是GPU的主要计算单元。由于SM承载了张量核心等专用加速器,其高效利用对训练效率至关重要。不幸的是,现有的集体通信系统与计算竞争SM资源,因为它们消耗SM用于与通信相关的数据移动和同步操作。我们观察到,通信原则上可以由DMA引擎驱动,从而消除通信中的SM参与。基于这一洞察,我们提出了HOCCL,一个零SM集体通信框架,由三个组件组成:流管理器、点对点(P2P)执行器和集体调度器。流管理器保持与其他GPU内核的操作级时间顺序。P2P执行器实现零SM点对点通信,而集体调度器编排P2P传输以最大化带宽。实验表明,HOCCL保持了接近峰值的通信性能,平均达到最先进水平的3%以内,同时消除了近10%总GPU SM上的通信占用。通过释放SM资源用于计算,HOCCL将端到端训练吞吐量提高了最多5%。

英文摘要

Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems compete with computation for SMs, as they consume SMs for communication-related data movement and synchronization operations. We observe that communication can, in principle, be driven by DMA engines, thereby eliminating SM involvement in communication. Based on this insight, we propose HOCCL, a zero-SM collective communication framework consisting of three components: a stream manager, a point-to-point (P2P) executor, and a collective scheduler. The stream manager preserves operator-level temporal ordering with other GPU kernels. The P2P executor enables zero-SM point-to-point communication, while the collective scheduler orchestrates P2P transfers to maximize bandwidth. Experiments show that HOCCL preserves near-peak communication performance, achieving within 3% of the state of the art on average, while eliminating communication occupancy on nearly 10% of total GPU SMs. By freeing SM resources for computation, HOCCL improves end-to-end training throughput by up to 5%.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • China Mobile (Suzhou) Software Technology Co., Ltd.(中国移动(苏州)软件技术有限公司)
  • Pengcheng Laboratory(鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑