AI 中文总结
研究高能物理实验中 Level-1 硬件神经网络触发器的故障,以实时处理大量数据。通过 RTL 故障注入研究 Belle II 触发系统中 GNN-ETM 的三种故障模式,发现现有验证基础设施的监测不对称,提出更准确的监测方法。
AI 中文摘要
随着粒子物理探测器规模扩大,高能物理实验需处理的数据量不断增加。在现场可编程门阵列上实现且越来越多地使用神经网络算法的一级触发系统实时过滤数据。但靠近交互点使其易受辐射影响,可能导致输出损坏、处理管道停滞或硬件损坏。本文首次对已部署的一级硬件神经网络触发器 Belle II 触发系统中的 GNN-ETM 进行寄存器传输级故障注入研究。针对对实时触发管道影响最大的三种故障模式:死锁、超时和数据包完整性违规。通过两项互补活动,在 211245 个信号上注入 1442840 个单粒子翻转。我们发现在现有验证基础设施中存在监测不对称,并提出阶段间活跃度监测作为仅输出观察的更准确替代方案,表明两种方法的平均无故障时间估计相差高达 78.7%。由此产生的每个阶段的数据确定了最高优先级的加固目标。
英文摘要
As particle physics detectors grow in scale, High Energy Physics experiments must process ever-increasing data volumes. Level-1 trigger systems, implemented on Field-Programmable Gate Arrays and increasingly using neural-network algorithms, filter this data in real time. However, their proximity to the interaction point exposes them to radiation, which can corrupt outputs, stall processing pipelines, or damage hardware, with significant financial and scientific consequences. In this work, we present the first Register Transfer Level fault-injection study of a deployed Level-1 hardware neural-network trigger, GNN-ETM in the Belle II trigger system. We target three failure modes most consequential to a real-time trigger pipeline: deadlocks, timeouts, and packet-integrity violations. Through two complementary campaigns, we inject 1 442 840 Single-Event Upsets across 211 245 signals. We find a monitoring asymmetry in the existing verification infrastructure and propose inter-stage liveness monitoring as a more accurate alternative to output-only observation, showing that Mean Time To Failure estimates from the two approaches differ by up to 78.7%. The resulting per-stage data identifies the highest-priority hardening targets.
CommentsAccepted to IEEE System On Chip Conference (SOCC) 2026