arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VARA:一种用于高能效计算的电压感知型基于RRAM的加速器

VARA: A Voltage-Aware ReRAM-Based Accelerator for Energy-Efficient Computing

Peng Dang, Yintao He, Huawei Li

arXiv 2609.00421首次发表:更新:

发表机构

SKLP, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences(中国科学院计算技术研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出电压感知型RRAM加速器VARA,通过电压感知训练算法提升激活稀疏性、共零激活重排方案实现交叉阵列级计算跳过,可大幅降低系统能耗并提升能效,且精度损失极小。

AI 中文摘要

基于RRAM的存内计算(IMC)架构被广泛视为缓解传统架构计算瓶颈的有前景方案。由于RRAM交叉阵列在模拟域执行矩阵-向量乘法(MVM),其计算能耗高度依赖权重与激活分布。然而,多数现有RRAM加速器主要聚焦于权重优化,却对激活对计算能耗的影响关注有限,使得激活稀疏性的节能潜力未得到充分挖掘。本文提出一种电压感知型基于RRAM的加速器(VARA)及其配套设计方法。具体而言,我们首先引入电压感知训练(VAT)算法,该算法在激活函数中加入预设阈值,以引导激活分布向零值靠拢,从而提升激活稀疏性。在此基础上,我们进一步提出交叉阵列级计算跳过的共零激活重排(CAR)方案,CAR基于激活维度的共零相关性对其进行聚类,并对激活矩阵及其对应权重进行一致重排。该过程将分散的零激活整合为连续的零值区域,以最大化交叉阵列级计算跳过的收益。大量实验结果表明,仅伴随极小的精度损失,与基线相比,VARA可将平均系统总能耗降低60.12%,并将平均系统能效提升2.68倍,优于现有针对稀疏激活优化的最先进加速器。

英文摘要

ReRAM-based in-memory computing (IMC) architectures are widely regarded as a promising approach to alleviating the computational bottleneck of conventional architectures. Since ReRAM crossbars perform matrix-vector multiplication (MVM) in the analog domain, their computational energy consumption is highly dependent on weight and activation distributions. However, most existing ReRAM accelerators focus primarily on weight optimization while paying limited attention to the impact of activations on computational energy consumption, leaving the energy-saving potential of activation sparsity largely underexploited. In this paper, we propose a voltage-aware ReRAM-based accelerator (VARA), along with its accompanying design methodology. Specifically, we first introduce a voltage-aware training (VAT) algorithm that incorporates a preset threshold into the activation function to steer the activation distribution toward zero values, thereby enhancing activation sparsity. Building upon this, we further propose a co-zero activation reordering (CAR) scheme for crossbar-level computation skipping. CAR clusters activation dimensions based on their co-zero correlations and consistently reorders both the activation matrix and its corresponding weights. This process consolidates scattered zero activations into contiguous zero-valued regions to maximize the benefits of crossbar-level computation skipping. Extensive experimental results demonstrate that, with only marginal accuracy loss, VARA reduces the average total system energy consumption by 60.12\% and improves the average system energy efficiency by 2.68$\times$ compared to the baseline, outperforming existing state-of-the-art accelerators for sparse-activation optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑