发表机构
Northeastern University; Microsoft Corporation(东北大学; 微软公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有损广域网多路径RDMA中数据包乱序和丢失问题,提出首个基于FPGA的完全硬件卸载数据包到达跟踪设计COMET,采用可扩展缓存架构,实现400Gbps线速并支持6倍并发连接。
AI 中文摘要
现代AI工作负载日益依赖于跨多个数据中心互连形成单一“AI工厂”的规模架构,以克服单个站点的功耗和散热限制。然而,将远程直接内存访问(RDMA)扩展到广域网(WAN)带来了根本性挑战:多路径数据包重排序、高延迟以及严重降低性能的数据包丢失。虽然纠删码(EC)已成为一种有前景的丢失恢复机制,但其有效性关键取决于硬件中实现的高效数据包到达跟踪。我们提出了紧凑多路径纠删码跟踪(COMET),这是首个在基于FPGA的网络接口卡(NIC)上实现的完全硬件卸载的数据包到达跟踪设计,用于有损广域网上的多路径RDMA。COMET采用可扩展的基于缓存的架构,支持在高链路速率下运行。我们的评估表明,COMET在400 Gbps及以上的速率下维持线速运行。关键的是,COMET将片上内存占用与链路带宽-延迟积(BDP)解耦,其基于缓存的架构(COMET缓存)能够支持比最先进的基于SoC的设计多6倍的并发连接。这些结果表明,可扩展的、完全硬件卸载的数据包到达跟踪在当前数据速率下的FPGA NIC上是可行的,其架构可扩展性延伸至新兴的1.6 Tbps NIC及更高速度。
英文摘要
Modern AI workloads increasingly rely on scale across architectures that interconnect multiple datacenters to form a single "AI factory", overcoming the power and cooling constraints of individual sites. However, extending Remote Direct Memory Access (RDMA) across wide area networks (WANs) introduces fundamental challenges: multi-path packet reordering, high latency, and packet loss that severely degrade performance. While erasure coding (EC) has emerged as a promising mechanism for loss recovery, its effectiveness critically depends on efficient packet arrival tracking implemented in hardware. We present COmpact Multi-path Erasure-coded Tracking (COMET), the first fully hardware-offloaded packet-arrival tracking design implemented on an FPGA-based network interface card (NIC) for multi-path RDMA over lossy WANs. COMET employs a scalable cache-based architecture that supports operation at high link rates. Our evaluation shows that COMET sustains line rate operation at 400 Gbps and beyond. Critically, COMET decouples on-chip memory footprint from link Bandwidth-Delay Product (BDP), and its cache-based architecture (COMET Cache) enables supporting 6 times more concurrent connections than state-of-the-art (SOTA) SoC-based designs. These results demonstrate that scalable, fully hardware-offloaded packet-arrival tracking is practical on FPGA-based NICs at current data rates, and its architectural scalability extends to emerging 1.6 Tbps NICs and beyond.
CommentsThis paper appeared in the 36th International Conference on Field-Programmable Logic and Applications https://2026.fpl.org/