arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StrataCL:面向生产级超级节点的 fabric 原生通信库

StrataCL: Fabric-Native Communication Library for Production Supernodes

Tiancheng Hu, Jin Qin, Yuzheng Wang, Ke Liu, TangShengsheng Li, Sheng Wang, Zhongzhe Hu, Tianlun Hu, Wei Wang, Lijun Li, Jingbin Zhou, Xiaoming Bao, Hongwei Sun, Jieru Zhao, Huimin Cui, Tao Xie, Chenxi Wang

arXiv 2607.26444首次发表:更新:

AI 中文总结

针对分布式 AI 工作负载的通信瓶颈,提出零冗余的 fabric 原生通信库 StrataCL,通过分配时注册等技术优化,在华为 CloudMatrix384 上显著提升了带宽与多类工作负载的性能。

AI 中文摘要

现代分布式 AI 工作负载运行在数百个加速器上,通信成为主要瓶颈。现有通信库大多以缓冲区为中心,因用户缓冲区与通信缓冲区分开管理,导致冗余数据拷贝或高昂的用户缓冲区注册开销。本文提出 StrataCL,一种零冗余的 fabric 原生通信库,面向生产级超级节点。StrataCL 引入「分配时注册」机制实现用户缓冲区直接通信,设计通信算子时采用工作负载均衡的 NPU 核心划分及 NPU 驱动的 SDMA 卸载,以充分利用超级节点架构特性。在华为 CloudMatrix384 平台上,StrataCL 使集体总线带宽提升最高 1.6 倍,MoE 调度/合并总线带宽提升最高 1.4 倍;在三个生产级工作负载上,StrataCL 使 LLM 推理吞吐量提升 1.9 倍,P99 TTFT 降低 2.2 倍,LLM 训练迭代时间缩短 1.4 倍,Recsys 训练迭代时间缩短 1.3 倍。

英文摘要

Modern distributed AI workloads run across hundreds of accelerators, making communication a major bottleneck. Existing communication libraries remain largely buffer-centric because user and communication buffers are managed separately, causing redundant data copies or costly user-buffer registration. This paper presents StrataCL, a zero-redundancy and fabric-native communication library for production supernodes. StrataCL introduces registration-on-allocation to realize user-buffer direct communication, and designs communication operators with workload-balanced NPU-core partitioning and NPU-driven SDMA offloading to exploit supernode architecture features. On the Huawei CloudMatrix384, StrataCL improves collective bus bandwidth by up to 1.6x and improves MoE dispatch/combine bus bandwidth by up to 1.4x. Across three production workloads, StrataCL improves LLM inference throughput by 1.9x, reduces P99 TTFT by 2.2x, and reduces LLM and Recsys training iteration time by 1.4x and 1.3x, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑