Zynq UltraScale+ MPSoC上开源机器学习加速器的质子辐照特性
Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoC
- Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg(卢森堡大学安全、可靠性与信任跨学科中心(SnT))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究为部署在Zynq UltraScale+ SoC上的开源Tensil NN加速器建立了质子辐照基准,揭示了其在特定辐照条件下的可用性损失与静默输出损坏,为空间用FPGA-SoC的软件加固提供了依据。
AI中文摘要:
随着星载计算系统越来越依赖神经网络(NN)加速器,商业黑盒架构的不透明性严重限制了可验证辐射缓解策略的开发。开源的、寄存器传输级(RTL)可访问的加速器通过支持用户自定义检测解决了这一限制,但很少有经验性的辐射响应基准。本研究为部署在Zynq UltraScale+ SoC上、执行ResNet-20推理的未受防护的开源Tensil NN加速器建立了基础的系统级质子辐照基准。在20至58 MeV质子辐照下,我们在受监控的操作窗口内注入了4.29×10^10 p/cm²的剂量。7次工作负载中断需要2次笔记本进程重启、4次重新启动或板卡重置,以及1次电源循环序列。2次输出损坏事件返回了不正确的CIFAR-10类别,但未丧失服务能力。在较长的事件中,加速器以正常节奏连续39次输入返回了CIFAR-10十种类别池之外的一个类别,进程保持存活,而内核日志、有限内存测试和采样功率均未显示异常,该卡住类别序列的观察在计划的比特流重新配置后结束。所有9次事件均发生在标称4 cm射束下,该射束暴露了SoC、LPDDR4和其他板卡电路;而在2 cm以SoC为中心的场下未发生任何事件。此模式显示出与场的关联,但未确定LPDDR4为原因,因为场大小与运行顺序和剂量存在混淆。Linux管理的加速器需要端到端内容检查和恢复,以达到损坏可能持续存在的状态。该基准记录了可用性损失和静默输出损坏,为空间系统中用于神经网络推理的商用现货FPGA-SoC的未来软件加固提供了支持。
英文摘要:
As spaceborne computing systems increasingly rely on neural network (NN) accelerators, the opacity of commercial, black-box architectures severely restricts the development of verifiable radiation mitigation strategies. Open-source, register-transfer level (RTL)-accessible accelerators resolve this limitation by enabling user-defined instrumentation, yet few have empirical radiation-response baselines. This work establishes a foundational system-level proton-irradiation baseline for an unmitigated open-source Tensil NN accelerator deployed on a Zynq UltraScale+ SoC executing ResNet-20 inference. Under 20 to 58 MeV proton irradiation, we delivered $4.29 \times 10^{10}$ p/cm$^{2}$ within monitored operational windows. Seven workload interruptions required two restarts of the notebook process, four reboots or board resets, and one power-cycle sequence. Two output-corruption events returned incorrect CIFAR-10 classes without loss of service. In the longer event, the accelerator returned a class absent from the ten-image CIFAR-10 pool for 39 consecutive inputs at normal cadence. The process remained alive, while the kernel log, limited memory test, and sampled power showed no anomaly. Observation of the stuck-class sequence ended with scheduled bitstream reconfiguration. All nine onsets occurred under the nominal 4 cm beam, which exposed the SoC, LPDDR4, and additional board circuitry; none occurred under the 2 cm SoC-centered field. This pattern shows a field association but does not establish LPDDR4 as the cause because field size was confounded with run order and dose. Linux-managed accelerators require end-to-end content checks and recovery that reaches the state in which corruption can persist. This baseline documents availability loss and silent output corruption, supporting future software hardening of COTS FPGA-SoCs for neural-network inference in space systems.