arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39978cs.DC

LatencyLab:基于DPDK的FPGA SmartNIC P4流水线时延测量框架

LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICs

Pavani Kuppili, Zhaoyang Han, Yicheng Qian, Suranga Handagala, Michael Zink, Miriam Leeser, Robert Ricci

首次发表
浏览论文内容

中文总结 AI 辅助

LatencyLab提出基于DPDK的FPGA P4流水线时延测量框架,利用双端口多播与TSC时间戳,无需PHC/PTP或时钟同步,在Alveo U280上实现纳秒级可重复测量,并验证了其准确性。

中文摘要 AI 辅助

P4可编程FPGA SmartNIC将数据包处理直接置于线速路径上,但开源的FPGA P4工具流未在流水线边界暴露时间戳功能,因此P4程序在目标FPGA上增加的时延很少被测量。本文提出LatencyLab,一个基于DPDK的FPGA P4流水线时延测量框架,该框架既不需要数据路径上的PHC/PTP支持,也不需要时钟同步。FPGA的两个端口共享同一网段,因此交换机将每个探测数据包的副本多播到两个端口:一个副本经过VitisNetP4流水线,另一个经过匹配的旁路路径。内核旁路DPDK接收器忙轮询两个端口,并在从NIC接收环形缓冲区取回每个数据包时,使用CPU时间戳计数器(TSC)为其打上时间戳。两个副本的到达时间差在针对承载相同流量的空比特流校准后,即可隔离出流水线时延;发送时间在减法中抵消。我们在AMD Alveo U280上评估了四个VitisNetP4程序,每个程序使用20000个数据包的轨迹进行探测,每个会话测量十次,共五个独立会话,所有数据包均经过TSC时间戳并反射用于硬件时间戳。测量得到的时延分布紧密且可重复:99%的数据包落在中位数20纳秒以内,会话中位数在1到2纳秒内重复(FiveTuple 107/137纳秒,Forward 149纳秒,RemoveHeader 177纳秒,Checksum 364纳秒,时钟频率250 MHz)。两个独立检查与框架结果一致:一个无内核反射器将每个探测对返回给ConnectX-5 NIC,其适配器时钟逐分位数地重现测量分布,误差在几纳秒内,且每个测量数据包比供应商的最坏情况时延界限低18到22个时钟周期。

英文摘要

P4-programmable FPGA SmartNICs place packet processing directly on the wire, but open FPGA P4 toolflows do not expose timestamping at the pipeline boundary, so the latency a P4 program adds on the target FPGA is rarely measured. This paper presents LatencyLab, a DPDK-based measurement framework for FPGA P4 pipeline latency that needs neither PHC/PTP support on the datapath nor clock synchronization. The FPGA's two ports share a network segment, so the switch multicasts a copy of each probe packet to both: one copy passes through the VitisNetP4 pipeline, the other through a matched bypass path. A kernel-bypass DPDK receiver busy-polls both ports and timestamps every packet with the CPU timestamp counter (TSC) as it is retrieved from the NIC's receive circular buffer. The arrival-time difference of the two copies isolates the pipeline latency after calibration against a null bitstream carrying the same traffic; transmit time cancel in the subtraction. We evaluate four VitisNetP4 programs on an AMD Alveo U280, probing each with a 20,000-packet trace measured ten times per session over five independent sessions, all TSC-timestamped and reflected for hardware timestamping. The measured latency distributions are tight and reproducible: 99% of packets fall within 20 ns of the median, session medians repeating within 1 to 2 ns (FiveTuple 107/137 ns, Forward 149 ns, RemoveHeader 177 ns, Checksum 364 ns at 250 MHz). Two independent checks agree with the framework: a kernel-free reflector returns every probe pair to a ConnectX-5 NIC whose adapter clock reproduces the measured distributions within a few nanoseconds, quantile by quantile, and every measured packet falls 18 to 22 clock cycles below the vendor's worst-case latency bound.

发表机构

  • University of Utah(犹他大学)
  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑