AI 中文总结
研究在FPGA SoC上实现容错的难题,提出分层数字孪生框架EIR,利用处理系统空闲周期承载互补孪生实现自主故障检测和恢复,实验表明其在故障覆盖率、能量和面积方面优于DMR基线,为弹性边缘AI部署提供实用途径。
AI 中文摘要
卷积神经网络(CNN)越来越多地部署在片上系统(SoC)平台上,硬件加速推理可实现低延迟边缘计算。在这些设备上实现容错仍然具有挑战性,因为传统冗余(双/三模冗余,DMR/TMR)会带来高昂的资源成本,而以软件为中心的方法(如基于算法的容错(ABFT)、检查点重启、指令级复制以及软件看门狗/断言)会引入显著的延迟/能量开销,降低模型精度,或对加速器引发的故障覆盖不足。本文提出了仿真完整性副本(EIR),这是一种用于FPGA SoC的分层数字孪生框架,可提供自主故障检测和恢复。与复制硬件逻辑并产生成比例的面积和功耗开销的DMR/TMR不同,EIR通过利用处理系统(PS)中的时间松弛来避免结构级复制。在可编程逻辑(PL)中加速器执行期间,PS通常未被充分利用;EIR利用这些空闲周期来承载两个互补的孪生:(i)兔子:用于快速故障检测的粗粒度行为模型;(ii)乌龟:用于从检查点状态进行精确恢复的细粒度门级模型。利用加速器的执行速度分析定期捕获加速器状态,以平衡性能开销和弹性。在代表性工作负载上的实验表明,相对于DMR基线,EIR在评估的故障模型和工作负载假设下实现了高经验故障覆盖率,同时降低了能量和面积,为在严格资源预算下实现弹性边缘人工智能部署指明了一条实用途径。
英文摘要
Convolutional neural networks (CNNs) are increasingly being deployed on system-on-chip (SoC) platforms, where hardware-accelerated inference enables low-latency edge computing. Achieving fault tolerance on these devices remains challenging because conventional redundancy (dual/triple modular redundancy, DMR/TMR) incurs high resource cost, while software-centric methods (e.g., algorithm-based fault tolerance (ABFT), checkpoint-restart, instruction-level duplication, and software watchdogs/assertions) introduce nontrivial latency/energy overheads, reduce model accuracy, or provide inadequate coverage for accelerator-induced faults. In this paper, we propose Emulated Integrity Replica (EIR), a hierarchical digital-twin framework for FPGA SoCs that provides autonomous fault detection and recovery. Unlike DMR/TMR, which replicates hardware logic and incurs proportional area and power overheads, EIR avoids fabric-level duplication by exploiting temporal slack in the processing system (PS). During accelerator execution in the programmable logic (PL), the PS typically remains underutilized; EIR capitalizes on these idle cycles to host two complementary twins: (i) Rabbit: a coarse-grained behavioral model for rapid fault detection and (ii) Tortoise: a fine-grained gate-level model that performs precise recovery from checkpointed states. The accelerator state is captured periodically, leveraging the accelerator's execution-speed profiling to balance performance overhead and resilience. Experiments on representative workloads show that EIR achieves high empirical fault coverage relative to a DMR baseline while reducing energy and area under the evaluated fault model and workload assumptions, indicating a practical path to resilient edge-AI deployments under strict resource budgets.
Comments10 Pages, 3 Figures, 3 Tables