arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

A-GHOST:面向可编程GPU推理的触发级数据高速流式传输

A-GHOST: High-rate streaming of trigger-level data to programmable GPU inference

I. Xiotidis, N. Clarke Hall, M. S. Larson, R. Gurunathan, C. Burdick, T. Du, N. Konstantinidis, K. Kordas, D. Leshchev, D. W. Miller, V. A. Petrovic, D. Sampsonidou, A. Thompson, T. Wengler

arXiv 2610.01761首次发表:更新:

AI 中文总结

A-GHOST提出将高能物理触发数据高速流式传输至GPU后端,利用NVIDIA IGX Thor实现40-100 Gbps吞吐和0.118毫秒延迟,验证了超越FPGA的复杂推理能力。

AI 中文摘要

A-GHOST(一种全局异构在线侦察触发器)是一项研发工作,旨在研究高能物理(HEP)实验中用于高速流式读取的架构。其核心思想是将靠近探测器的确定性前端和固定延迟处理从决策平台转变为聚合和流式数据源。紧凑的触发级数据被流式传输到支持GPU的后端,在那里可以运行比硬件触发器资源和延迟范围更复杂的算法。本文展示了使用NVIDIA IGX Thor开发套件的概念验证后端。一个软件速率控制的发送器通过QSFP接口和外部环回电缆将数据发送到第二个QSFP接口,从而可以在与FPGA源集成之前研究网络到GPU的路径。NVIDIA DAQIRI支持直接接收数据到GPU可访问内存中,而自定义CUDA内核将数据包负载重新组装成持久、连续的TensorRT输入窗口,无需数据类型转换,该转换由模型处理。使用源自HEP的事件表示(由十个主要量热器簇组成),后端在保持恒定输入速率和p95间隔推理延迟为0.118毫秒的情况下,从40 Gbps扩展到100 Gbps。系统连续运行数小时,无数据丢失或推理错误。生成式和时序感知神经网络展示了超出通常部署在固定延迟FPGA触发系统中的工作负载。该结果确立了A-GHOST作为未来FPGA到GPU流式读取系统的基础。

英文摘要

A-GHOST (A Global Heterogeneous Online Scouting Trigger) is an R&D effort investigating high-rate streaming readout architectures for high-energy physics (HEP) experiments. The central idea is to convert the deterministic front-end and fixed-latency processing close to the detector from a decision-making platform to an aggregation and streaming source. Compact trigger-level data are streamed to a GPU-enabled backend, where substantially more complex algorithms can run beyond the resource and latency envelope of the hardware trigger. This paper presents a proof-of-concept backend using the NVIDIA IGX Thor development kit. A software rate-controlled transmitter sends data through a QSFP interface and external loopback cable to a second QSFP interface, allowing the network-to-GPU path to be studied before integration with FPGA sources. NVIDIA DAQIRI enables reception directly into GPU-accessible memory, while a custom CUDA kernel reassembles packet payloads into persistent, contiguous TensorRT input windows without data-type conversion, which is handled by the models. Using a HEP-derived event representation consisting of the ten leading calorimeter clusters, the backend scales from 40 to 100 Gbps while sustaining a constant input rate and an inference latency of 0.118 ms at the p95 interval. The system operates for multiple hours without drops or inference errors. Generative and temporally aware neural networks demonstrate workloads beyond those normally deployed in fixed-latency FPGA trigger systems. The result establishes A-GHOST as a basis for future FPGA-to-GPU streaming readout systems.

Comments17 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑