arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25560cs.DC

Co-Fabric:打破主机域边界,实现统一的xPU互连

Co-Fabric: Breaking Host-Domain Boundaries for Unified xPU Interconnection

  • IEIT SYSTEMS Co., Ltd.(易华录系统有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Zhen Peng, Jiaming Huang, Chaofan Chen, Zhao Zhang, An Wu, Baoyang Liu, Xinglong Wang, Tanlong Ci, Jinfeng Li, Xueke Duan, Hao Wang, Xi Chen, Shunshun Zhang, Zh… 展开作者

Zhen Peng, Jiaming Huang, Chaofan Chen, Zhao Zhang, An Wu, Baoyang Liu, Xinglong Wang, Tanlong Ci, Jinfeng Li, Xueke Duan, Hao Wang, Xi Chen, Shunshun Zhang, Zhiyuan Su, Zhu Cao, Zhichong Dou, Shaohua Wu, Lu Jing, Yue Yuan

AI总结:

Co-Fabric提出一种打破主机域边界的总线式互连,通过精简四层协议栈、跨域P2P路由和统一地址空间,在64-xPU系统上相比RoCE降低延迟超50%、提升带宽2-5倍,并显著降低成本和功耗。

AI中文摘要:

大模型参数已超出单个xPU的容量,分散在跨越不同主机域的多个xPU上,其中xPU到xPU的通信主导着整体系统效率。现有的纵向扩展互连仍显不足:基于以太网的网络解决方案(如RoCE(融合以太网上的RDMA))引入了特定的消息语义和协议栈特性,并依赖于碎片化的每主机寻址,而传统基于主机的结构局限于单个主机域,缺乏跨主机的统一寻址。本文提出了Co-Fabric,一种基于总线的互连,与常规总线设计不同,它打破了主机域边界,为纵向扩展超级pod提供统一的xPU互连。Co-Fabric做出三项贡献:一个精简的四层协议栈,实现纳秒级处理延迟并具备原生可靠性;一种跨域扩展和P2P机制,利用数据包头中嵌入的端口标识符进行路由;以及一个基于影子设备自动枚举的统一地址空间。在64-xPU 3D-Mesh系统上,Co-Fabric将节点间通信延迟降低超过50%,带宽比RoCE提升2-5倍,使DeepSeek R1推理加速30%-80%。此外,由于其精简的四层协议栈和更高的数据通信效率减少了相对于基于以太网的RoCE栈的协议和处理开销,Co-Fabric将互连本身的成本和功耗分别降低高达80%和5%。这些结果展示了Co-Fabric在AI计算中心的优势。

英文摘要:

Large-model parameters have grown beyond the capacity of a single xPU, dispersing across multiple xPUs spanning distinct host domains, where xPU-to-xPU communication dominates overall system efficiency. Existing scale-up interconnect remains inadequate: network-based solutions built on Ethernet--such as RoCE (RDMA over Converged Ethernet)--introduce specific message-semantics and protocol-stack characteristics, and rely on fragmented per-host addressing, while conventional host-based fabrics are confined to a single host domain and lack cross-host unified addressing. This paper presents Co-Fabric, a bus-based interconnect that, unlike conventional bus designs, breaks host-domain boundaries to deliver unified xPU interconnection for scale-up superpods. Co-Fabric makes three contributions: a streamlined four-layer protocol stack achieving nanosecond-scale processing latency with native reliability; a cross-domain scaling and P2P mechanism that routes using port identifiers embedded in the packet header; and a unified address space built on shadow-device auto-enumeration. On a 64-xPU 3D-Mesh system, Co-Fabric cuts inter-node communication latency by over 50% and improves bandwidth by 2-5x over RoCE, accelerating DeepSeek R1 inference by 30%-80%. Moreover, since its streamlined four-layer protocol stack and higher data-communication efficiency reduce protocol and processing overhead relative to the Ethernet-based RoCE stack, Co-Fabric cuts the cost and power of the interconnect itself by up to 80% and 5%, respectively. These results demonstrate Co-Fabric's advantage for AI computing centers.

↑