arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HCCL:面向元训练与推理加速器的集合通信库

HCCL: Collective Communication for Meta Training and Inference Accelerators

Wesley Bland, Tiago Antunes, Lars Paul Huse, Chidambaram Muthu, Adel Abouchaev, Rabib Alam, Abdullah Alperen, Alexey Andronov, Jose Anto Akkara, Vineet Badhwar, Pavan Balaji, Daniel Berkovitch, Bartosz Bogdanski, Shmeelok Chakraborty, Sungjun Cho, John Choi, James Custer, Rodrigo De Castro, Nguyen Dinh Pham, Matthew Edwards, Kristian Evensen, Evan Ezell, Alex Finestead, Seth Goldstein, Prankur Gupta, Ranwei Hu, Adam Incera, Anand Jayaraman, Prashanth Kannan, Soumil Kanwal, Martin Karp, Sameer Kumar, Naina Kuruballi Mahesh, Wei Lin Guay, Cristian Lumezanu, Cory Modlin, Dag Georg Moxnes, Hoang Nam Nguyen, Ashay Narsale, Jaden Padua, Kirtesh Patil, Minh Pham, Amin Qassoud, Ashwin Ramachandran, David Ramon Prados, Pallavi Shurpali, Gregory R. Steinbrecher, John Sundharam, Vangelis Tasoulas, Fuhou Tian, Srinivas Vaidyanathan, Vimal Vasudevan, Nicolaas Viljoen, Daniel Winkelman, Yijing Zeng, Zhaoqi Zhu, Stig Arne Olsen, Gilad Goldfarb, Rajiv Krishnamurthy, Rajeev Nair, Jonas Olsson, Joseph Provine, Sreeram Ravinoothala, Shivayogi Ugaji, Hongyi Zeng, Nairan Zhang

arXiv 2608.00358首次发表:更新:

AI 中文总结

该研究提出与Meta MTIA 300加速器协同设计的HCCL集合通信库,通过编译式通信模型等优化,在训练和推理场景下提升通信性能并降低延迟,同时减少对计算吞吐量的影响。

AI 中文摘要

我们提出HCCL,这是与Meta的MTIA 300加速器协同设计的集合通信库,MTIA 300是首款在芯片封装上直接集成后端网络的Meta芯片。MTIA 300配备了近内存计算(NMC)的专用消息引擎(ME),可将集合操作完全从计算网格中卸载,实现计算与通信的高度重叠。HCCL采用编译式通信模型,主机生成包含依赖关系的每个集合操作的完整描述。我们阐述了控制与数据路径架构、针对MTIA 300非对称 scale-up 和 scale-out 网络的拓扑感知算法选择,以及针对训练和推理工作负载的优化。对于训练,HCCL在机架内集合操作上实现高达940 GB/s的吞吐量,同时对并发计算吞吐量的影响低于0.5%;对于推理,我们利用绕过调度路径的单边原语以最小化集合延迟,并阐述了针对延迟敏感工作负载优化计算-通信流水线的集合设计。

英文摘要

We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package. MTIA 300 includes dedicated message engines (MEs) with near-memory compute (NMC) that fully offload collective execution from the compute grid, enabling large overlap between computation and communication. HCCL uses a compiled communication model in which the host generates a complete description of each collective including dependencies. We describe the control and data path architecture, topology-aware algorithm selection across MTIA 300's asymmetric scale-up and scale-out network, and optimizations for both training and inference workloads. For training, HCCL achieves up to 940 GB/s on intra-rack collectives while introducing less than 0.5% degradation to concurrent compute throughput. For inference, we leverage one-sided communication primitives that bypass the scheduling path to minimize collective latency and describe collective designs that improve compute-communication pipelining for latency-sensitive workloads.

Comments12 pages, 17 figures, to be published in the proceedings of "SC '26: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis"

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑