迈向复杂SmartNIC的简单模型
Towards Simple Models of Complex SmartNICs
浏览论文内容
中文总结 AI 辅助
本文提出ZRAM模型,通过平台图与程序图映射及三个指标(能力、屋顶线、容量)评估SmartNIC应用放置的可行性与瓶颈,并提炼七个设计模式,为复杂SmartNIC编程提供简单预测方法。
中文摘要 AI 辅助
云供应商将雄心勃勃的网络内处理(例如加密、遥测)推向网卡(NIC)以卸载服务器,即使链路速率攀升至太比特速度。供应商以异构SmartNIC作为回应。例如,NVIDIA BlueField-3在线上(wire)和主机CPU之间插入了线速eSwitch、多线程数据路径加速器(Data-Path Accelerator)、通用ARM核心以及大量固定功能加速器。这些设备以难以编程著称,更难预测性能。应用程序可以通过多种方式实现,每种选择都可能遇到不同的瓶颈。设计者理想情况下需要——在编写任何代码之前,以低成本的方式——了解可行的选择及其瓶颈,以及改进性能的设计模式。我们的论文提供了一个起点,使用我们称之为ZRAM的模型来回答这些问题。该模型将处理区域及其通道的平台图与任务及其流量比例的程序图配对。应用程序设计者或编译器选择一个放置方案,将程序图映射到平台图上。直接从该映射计算出的三个指标——能力(capability)、屋顶线(roofline)和容量(capacity)——对放置方案进行评分,决定其可行性并指出瓶颈资源。我们以DDoS检测器作为主要案例研究,并简要探讨了另外两个应用:决策树推理和RDMA遍历。我们提炼了七个编程SmartNIC的设计模式,其中包括一个关键模式,我们称之为“筛选”(sifting)。ZRAM可推广到其他SmartNIC,如Intel IPU E2200和AMD Pensando Salina 400,并开辟了一个新的研究议程,包括编译器和硬件设计。
英文摘要
Cloud vendors push ambitious in-network processing (e.g., crypto, telemetry) onto the NIC to offload servers even as link rates climb to terabit speeds. Vendors have responded with heterogeneous SmartNICs. For example, NVIDIA BlueField-3 interposes---between the wire and the host CPUs---a line-rate eSwitch, a multithreaded Data-Path Accelerator, general-purpose ARM cores, and a sea of fixed-function accelerators. These devices are notoriously hard to program, and harder still to predict. Applications can be implemented in many ways, with each choice potentially hitting a different bottleneck. A designer ideally needs to know---cheaply, and before a line of code is written---feasible choices and their bottlenecks, and design patterns to improve performance. Our paper offers a starting point to answer these questions using what we call the ZRAM model. It pairs a platform graph of processing zones and their channels with a program graph of tasks and their traffic fractions. The application designer or a compiler chooses a placement that maps the program graph onto the platform graph. Three metrics computed directly from this mapping---capability, roofline, and capacity---score the placement, deciding its feasibility and naming the bottleneck resource. We use a DDoS detector as a primary case study, and briefly explore two other applications, decision-tree inference and RDMA traversal. We distill seven design patterns for programming SmartNICs including a key one we call sifting. ZRAM generalizes to other SmartNICs such as Intel IPU E2200 and AMD Pensando Salina 400, and opens a new research agenda that includes compilers and hardware design.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。