AI 中文总结
该研究提出与Meta MTIA 300加速器协同设计的HCCL集合通信库,通过编译式通信模型等优化,在训练和推理场景下提升通信性能并降低延迟,同时减少对计算吞吐量的影响。
AI 中文摘要
我们提出HCCL,这是与Meta的MTIA 300加速器协同设计的集合通信库,MTIA 300是首款在芯片封装上直接集成后端网络的Meta芯片。MTIA 300配备了近内存计算(NMC)的专用消息引擎(ME),可将集合操作完全从计算网格中卸载,实现计算与通信的高度重叠。HCCL采用编译式通信模型,主机生成包含依赖关系的每个集合操作的完整描述。我们阐述了控制与数据路径架构、针对MTIA 300非对称 scale-up 和 scale-out 网络的拓扑感知算法选择,以及针对训练和推理工作负载的优化。对于训练,HCCL在机架内集合操作上实现高达940 GB/s的吞吐量,同时对并发计算吞吐量的影响低于0.5%;对于推理,我们利用绕过调度路径的单边原语以最小化集合延迟,并阐述了针对延迟敏感工作负载优化计算-通信流水线的集合设计。
英文摘要
We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package. MTIA 300 includes dedicated message engines (MEs) with near-memory compute (NMC) that fully offload collective execution from the compute grid, enabling large overlap between computation and communication. HCCL uses a compiled communication model in which the host generates a complete description of each collective including dependencies. We describe the control and data path architecture, topology-aware algorithm selection across MTIA 300's asymmetric scale-up and scale-out network, and optimizations for both training and inference workloads. For training, HCCL achieves up to 940 GB/s on intra-rack collectives while introducing less than 0.5% degradation to concurrent compute throughput. For inference, we leverage one-sided communication primitives that bypass the scheduling path to minimize collective latency and describe collective designs that improve compute-communication pipelining for latency-sensitive workloads.
Comments12 pages, 17 figures, to be published in the proceedings of "SC '26: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis"