AI 中文总结
本文提出Lonic,通过算法-硬件协同设计实现INT4精度的全局部在线SNN训练,在能效、速度等核心指标上显著优于Apple M4、Nvidia V100 GPU及多种专用加速器。
AI 中文摘要
脉冲神经网络(SNN)作为一种高能效学习范式,近年来受到越来越多的关注。现有研究也提出了时间维度上的全局部在线SNN训练算法,以解决内存和计算开销问题。然而,这些研究并未考虑算法的优势能否有效转化为实际设备的能效。为应对这一挑战,本文提出Lonic,这是一种面向高能效、可扩展全局部在线监督SNN学习的算法-硬件协同设计方案。在算法层面,本文实现了用于全局部在线SNN学习的INT4低精度训练算法,同时保持了模型精度。在硬件层面,为充分发挥所提算法的优势,本文引入了可重构无乘法器整数PE阵列、双优化零门控策略、时间前缀加速的局部学习数据流以及低精度权重迁移技术,以显著提升训练效率。与Apple M4和Nvidia V100 GPU相比,Lonic的平均能效分别提升17.44倍和66.28倍,加速比分别为3.25倍和1.02倍;此外,与类ASIC TPU和H2Learn加速器相比,Lonic的能效分别提升15.95倍(14.64倍),面积效率分别提升1.52倍(7.28倍)。Lonic的代码可在指定网址获取。
英文摘要
Spiking neural networks (SNNs) have recently attracted increasing attention as an energy-efficient learning paradigm. Existing works also propose temporally and fully local online SNN training algorithms to address memory and computation overhead. However, they do not consider whether the algorithmic advantages can be effectively translated into real-device efficiency. To address this challenge, we present Lonic, an algorithm-hardware co-design for energy-efficient and scalable fully local online supervised SNN learning. On the algorithm side, we implement an INT4 low-precision training algorithm for fully local online SNN learning while maintaining accuracy. On the hardware side, to leverage the benefits of the proposed algorithm, we introduce reconfigurable multiplier-free integer PE arrays, dual-optimization zero-gating strategy, temporal prefix-accelerated local learning dataflow, and low-precision weight movement to significantly improve training efficiency. Compared to Apple M4 and Nvidia V100 GPUs, Lonic achieves average energy efficiency improvements of 17.44x and 66.28x, respectively, along with speedups of 3.25x and 1.02x, respectively. Moreover, Lonic achieves 15.95x (14.64x) and 1.52x (7.28x) energy efficiency (area efficiency) over ASIC TPU-like and H2Learn accelerators, respectively. The code for Lonic is available at https://github.com/peilin-chen/Lonic.
CommentsAccepted to ICCAD 2026