arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

适用于内存深度学习加速器的抗更新干扰模拟ReRAM交叉阵列

Update Disturbance-Resilient Analog ReRAM Crossbar Arrays for In-Memory Deep Learning Accelerators

Wooseok Choi, Tommaso Stecconi, Donato Francesco Falcone, Matteo Galetta, Victoria Clerico, Elisa Zaccaria, Mamidala Saketh Ram, Antonio La Porta, Folkert Horst, Daniel Jubin, Matias Senger, Marilyne Sousa, Steffen Reidt, Ralph Heller, Bernabe Linares-Barranco, Valeria Bragaglia, Bert Jan Offrein

arXiv 2608.25781首次发表:更新:

AI 中文总结

本研究针对模拟ReRAM交叉阵列内存内训练的权重更新干扰难题,开发了350nm工艺的抗干扰ReRAM器件,验证了其在内存深度学习加速器全并行权重更新中的应用潜力。

AI 中文摘要

采用交叉阵列架构的阻变存储器(ReRAM)技术在模拟AI加速器硬件中具有巨大潜力,可实现内存内推理与训练。近期进展已通过将计算密集型训练工作负载卸载至片外数字处理器,成功实现推理加速,但训练算法的内存内加速对于构建更具可持续性和能效的AI至关重要,目前仍处于研究早期阶段。本研究聚焦于模拟ReRAM阵列的内存内训练加速,解决全并行权重更新过程中的关键挑战:交叉点器件的权重值干扰问题。提出了基于350nm硅工艺的ReRAM器件解决方案,利用HfOx层内的纳米级导电丝形成的阻变导电金属氧化物(CMO)。该器件不仅具备60纳秒的快速非易失性模拟开关特性,还展现出出色的抗更新干扰能力,可承受超过10万次脉冲。通过COMSOL Multiphysics仿真分析了该ReRAM的抗干扰性能,对导电丝诱导的热电能量集中进行建模,该现象导致器件对输入电压幅度呈现高度非线性响应。还在后端集成的ReRAM阵列芯片上展示了无干扰的并行权重映射。最后,全面的硬件感知神经网络仿真验证了本研究的ReRAM在具备全并行权重更新能力的内存深度学习加速器中的应用潜力。

英文摘要

Resistive memory (ReRAM) technologies with crossbar array architectures hold significant potential for analog AI accelerator hardware, enabling both in-memory inference and training. Recent developments have successfully demonstrated inference acceleration by offloading compute-heavy training workloads to off-chip digital processors. However, in-memory acceleration of training algorithms is crucial for more sustainable and power-efficient AI, but still in an early stage of research. This study addresses in-memory training acceleration using analog ReRAM arrays, focusing on a key challenge during fully parallel weight updates: disturbances of the weight values in cross-point devices. A ReRAM device solution is presented on 350 nm silicon technology, utilizing a resistive switching conductive metal oxide (CMO) formed on a nanoscale conductive filament within a HfOx layer. The devices not only exhibit 60 ns fast, non-volatile analog switching, but also demonstrates outstanding resilience to update disturbances, enduring over 100k pulses. The disturbance tolerance of the ReRAM is analyzed using COMSOL Multiphysics simulations, modeling the filament-induced thermoelectric energy concentration that results in a highly nonlinear device responses to input voltage amplitudes. Disturbance-free parallel weight mapping is also demonstrated on the back-end-of-line integrated ReRAM array chip. Finally, comprehensive hardware-aware neural network simulations validate the potential of our ReRAM for in-memory deep learning accelerators capable of fully parallel weight updates.

Journal refAdv. Sci. 13, no. 4 (2026): e04578

DOI:10.1002/advs.202504578

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑