arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图神经网络整数量化用于实时FPGA径迹查找

Integer Quantization of Graph Neural Networks for Real-Time FPGA Track Finding

Andrea Cardini, Pelayo Leguina, Elena Aller, Santiago Folgueras

arXiv 2609.28144首次发表:更新:

发表机构

Universidad de Oviedo; Instituto de Ciencias y Tecnológicas Espaciales de Asturias (ICTEA)(奥维耶多大学; 阿斯图里亚斯空间科技研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种可复现的位精确工作流程,通过训练后量化将图神经网络映射到FPGA,实现固定延迟推理,在CMS一级触发器中以19时钟周期延迟和75%准确率完成径迹查找。

AI 中文摘要

在CMS一级触发器中,针对位移μ子签名的实时径迹查找必须在12.5 μs的严格固定延迟约束下运行,同时处理高通量探测器数据。由于μ子击中自然映射到稀疏、不规则的图上,图神经网络(GNNs)成为有吸引力的候选方案;然而,将消息传递模型映射到现场可编程门阵列(FPGAs)需要数值精度、微架构和高层次综合(HLS)实现的协同设计。本工作提出了一种可复现、位精确的工作流程,桥接GNN设计与FPGA原型验证,用于固定延迟推理。该方法通过在两层的GraphSAGE网络上实现到XCVU13P FPGA上得到验证,使用Cora引文网络作为固定大小的基准进行固件评估,将部署可行性与物理任务解耦。从超出可用FPGA资源预算的32位浮点参考开始,我们通过训练后量化(INT8权重和激活,INT32偏置)、用算术移位替代缩放乘法器的2的幂次尺度近似,以及数据驱动的位宽缩减,推导出仅整数数据通路。每个阶段都在Vitis HLS C仿真中针对Python整数模拟器进行位精确验证。优化后的INT8 2的幂次设计在20% DSP、6% FF和27% LUT利用率下实现了19个时钟周期(标称360 MHz时钟下52.8 ns)的推理延迟,准确率为75.0±1.1%,而FP32为78.0±0.8%。准确率报告为多个训练种子的平均值。所得到的工作流程为CMS一级触发器中基于GNN的固定延迟径迹重建建立了一条具体、可迁移的路径。

英文摘要

Real-time track finding for displaced-muon signatures in the CMS Level-1 trigger must operate under strict fixed-latency constraints of 12.5 $μ$s while processing high-throughput detector data. Because muon hits map naturally onto sparse, irregular graphs, graph neural networks (GNNs) are attractive candidates; however, mapping message-passing models to field-programmable gate arrays (FPGAs) requires careful co-design of numerical precision, microarchitecture, and high-level synthesis (HLS) implementation. This work presents a reproducible, bit-exact workflow bridging GNN design and FPGA prototyping for fixed-latency inference. The methodology is demonstrated by implementing a two-layer GraphSAGE network onto an XCVU13P FPGA, using the Cora citation network as a fixed-size benchmark for firmware evaluation that decouples deployment feasibility from the physics task. Starting from a 32-bit floating-point reference that exceeds the available FPGA resource budget, we derive an integer-only datapath through post-training quantization (INT8 weights and activations, INT32 biases), a power-of-two scale approximation that replaces rescaling multipliers with arithmetic shifts, and data-driven bit-width narrowing. Every stage is validated bit-exactly against a Python integer emulator in Vitis HLS C-simulation. The optimized INT8 power-of-two design achieves an inference latency of 19 clock cycles (52.8 ns at the nominal 360 MHz clock) at 20\% DSP, 6\% FF, and 27\% LUT utilization, with 75.0$\pm$1.1\% accuracy compared to the 78.0$\pm$0.8\% for FP32. Accuracy is reported as the average across multiple training seeds. The resulting workflow establishes a concrete, transferable path toward fixed-latency GNN-based track reconstruction in the CMS Level-1 trigger.

CommentsThis work has been submitted to the IEEE TNS journal for possible publication. The work was originally presented at the 25th IEEE Real Time Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑