发表机构
Amirkabir University of Technology (Tehran Polytechnic)(阿米尔卡比尔理工大学(德黑兰理工学院))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对模拟ReRAM存内计算加速器的固定故障问题,提出基于冗余硬件与故障感知映射的容错方案,通过多目标优化确定冗余配置,恢复推理精度约22.39%,MTTF提升约61倍。
AI 中文摘要
模拟ReRAM存内计算(PIM)加速器为深度卷积神经网络(CNN)推理提供高并行性和高能效。然而,它们对永久性故障的敏感性,如高阻固定(SaH)和低阻固定(SaL)电阻状态,通过永久损坏映射到ReRAM单元电导值的CNN权重并降低推理精度,构成重大挑战,这导致安全关键应用中系统不可靠。在本文中,我们提出一种针对模拟ReRAM存内计算加速器的容错方案,以最小冗余开销解决固定故障(SAFs)并恢复分类精度下降。所提方案包含基于冗余的硬件解决方案以及故障感知映射方法,以确保ReRAM交叉阵列中可靠的模拟计算。我们分析了不同冗余行和列数量对精度和设计指标的影响。随后,构建并求解一个多目标优化(MOO)问题,以在考虑各种设计指标间权衡的情况下高效确定冗余行和列的数量。此外,针对双交叉阵列结构提出一种故障感知权重映射,以进一步补偿由SAFs引起的精度下降。仿真结果表明,对于使用MNIST数据集的SimpleNet模型,在四种最优解配置下,推理精度平均恢复约22.39%,每种配置在可靠性与面积、能耗和延迟开销之间提供权衡。与基线相比,平均无故障时间(MTTF)平均提升约61倍。与仅行和仅列配置相比,这些选定配置还平均降低能耗和面积开销32%。
英文摘要
Analog ReRAM-based process-in-memory (PIM) accelerators provide high parallelism and energy efficiency for deep convolutional neural networks (CNNs) inference. However, their susceptibility to permanent faults, such as stuck-at high (SaH) and stuck-at low (SaL) resistance states, poses a major challenge by permanently corrupting the CNN weights mapped to conductance values of ReRAM cells and degrading inference accuracy, which leads to system unreliability in safety-critical applications. In this paper, we propose a fault-tolerant scheme for analog ReRAM-based PIM accelerators to tackle stuck-at faults (SAFs) with minimal redundancy overhead to recover classification accuracy degradation. The proposed scheme contains a redundancy-based hardware solution alongside fault-aware mapping method for ensuring reliable analog computation in ReRAM crossbar. We analyze the impact of varying number of redundant rows and columns on accuracy and design metrics. Subsequently, a multi-objective optimization (MOO) problem is formulated and solved to efficiently determine the number of redundant rows and columns, considering trade-offs among various design metrics. Furthermore, a fault-aware weight mapping is proposed for dual-crossbar structures to further compensate for the accuracy degradation caused by SAFs. Simulation results show that, for the SimpleNet model using the MNIST dataset, the inference accuracy is recovered by approximately 22.39%, on average, across four configurations of optimal solutions, each offering a trade-off between reliability and area, energy consumption, and latency overheads. The mean-time-tofailure (MTTF) improves by about 61x on average compared to the baseline. These selected configurations also reduce energy and area overheads by 32%, on average, in comparison to row-only and column-only configurations.
Journal refDorostkar, Aniseh, Hamed Farbeh, and Hamid R. Zarandi. "ReSAFT: An efficient stuck-at fault-tolerant scheme for ReRAM-based process-in-memory accelerators." Future Generation Computer Systems 185 (2026): 108685
DOI:10.1016/j.future.2026.108685