AI 中文总结
研究事件驱动神经形态推理中权重稀疏性带来的问题及权衡,通过比较三种加速器量化稀疏性代价,利用特定流程和估计评估不同剪枝水平下的神经网络推理,结果表明性能和能量依赖稀疏性处理方式,SIMT在高稀疏性下有优势。
AI 中文摘要
事件驱动的神经形态推理通过仅在脉冲时更新神经元状态来利用激活稀疏性。然而,权重稀疏性会引入不规则的收集式更新,破坏锁步单指令多数据(SIMD)执行。我们将控制、元数据和内存活动中产生的开销称为稀疏性代价。本文通过比较集成在一个神经形态核心中的三种密切相关的加速器来量化这种代价:(i)基线锁步SIMD,(ii)位图门控的稀疏SIMD,它选择性地禁用通道而不压缩权重,以及(iii)具有每个处理元素地址生成和游程编码稀疏权重的单指令多线程(SIMT)风格设计。使用GF22FDX+中的RTL到门流程和活动驱动的能量估计,我们评估了不同训练后剪枝水平下的事件驱动神经网络推理。结果表明,由于SRAM占主导地位,核心总面积变化不显著,而性能和能量强烈依赖于稀疏性的处理方式:SIMD和稀疏SIMD表现出近乎恒定的吞吐量,稀疏SIMD由于位图和密集存储开销而实现有限的节能,SIMT在高稀疏性下提供最强的能量缩放和显著的加速,尽管由于元数据读取、负载不平衡和与稀疏性无关的阶段而具有亚线性增益。所提出的架构和实验的硬件代码可供研究目的公开访问。
英文摘要
Event-driven neuromorphic inference exploits activation sparsity by updating neuron state only on spikes. However, weight sparsity introduces irregular gather-style updates that undermine lockstep Single Instruction Multiple Data (SIMD) execution. We call the resulting overheads in control, metadata, and memory activity the sparsity tax. This paper quantifies that tax by comparing three closely related accelerators integrated into one neuromorphic core: (i) baseline lockstep SIMD, (ii) bitmap-gated Sparse-SIMD that selectively disables lanes without compressing weights, and (iii) a Single Instruction Multiple Threads (SIMT) style design with per-PE address generation and run-length coded sparse weights. Using an RTL-to-gates flow in GF22FDX+ and activity-driven energy estimation, we evaluate event-driven neural network inference across varying post-training pruning levels. Results show that total core area changes are insignificant because SRAM dominates area, while performance and energy strongly depend on how sparsity is handled: SIMD and Sparse-SIMD exhibit near-constant throughput, Sparse-SIMD achieves limited energy savings due to bitmap and dense-storage overheads, and SIMT provides the strongest energy scaling and substantial speedups at high sparsity, albeit with sublinear gains due to metadata reads, load imbalance, and sparsity-independent phases. The hardware code for the proposed architectures and experiments is publicly accessible for research purposes.
CommentsAccepted for publication in the IEEE MCSoC 2026 conference