发表机构
Northeastern University London(伦敦东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对脉冲Transformer零阶微调在内存计算加速器上的部署挑战,提出事件触发式隐式扰动架构,通过优化扰动生成方案降低硬件开销与能量,在准确率相当的同时大幅减少扰动相关能量。
AI 中文摘要
零阶(ZO)优化仅通过前向传播评估估计梯度,适用于对不可微、事件驱动的脉冲神经网络(SNN)进行微调。然而,其在内存计算(IMC)加速器上的部署受到两个问题限制:一是显式权重扰动产生的重复读-改-写(RMW)操作,二是用于每个权重统计独立扰动的随机数生成器(RNG)的硬件开销过高。为解决这些挑战,我们提出了隐式扰动零阶(IPZO)架构,该架构中由事件触发式扰动生成单元(PGU)计算的扰动和与IMC阵列产生的加权和相结合,消除了扰动诱导的RMW操作,同时保留了IMC的权重静态执行特性。通过利用脉冲稀疏性,PGU仅为脉冲激活的权重行生成并累加扰动贡献,从而减小了RNG阵列所需的行维度。我们进一步引入了地址驱动的异或重组方案(PGU-XOR),以缓解直接复用RNG(PGU-Reuse)导致的空间相关性问题。实验结果表明:(1)在Spikingformer/CIFAR-10上,PGU-XOR的准确率与软件RNG相当(76.41% vs. 76.53%),在SpikeGPT/WikiText-2上的困惑度(PPL)也相近(54.20 vs. 53.23);而PGU-Reuse的准确率下降9.56个百分点,PPL升高11.8;(2)采用台积电16纳米CMOS工艺实现时,相对于PGU-Reuse,PGU-XOR每次矩阵-向量乘法的面积开销为40.3%-46.0%,能量开销为15.2%-48.9%,但在相同准确率下,其更快的收敛速度将总扰动能量降至PGU-Reuse的0.51倍;(3)对于批量大小B=64、时间步长T=4的情况,IPZO将扰动能量降至传统显式权重扰动的0.46倍-0.83倍,且优势随BT的减小而增大。
英文摘要
Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for fine-tuning non-differentiable, event-driven spiking neural networks (SNNs). However, its deployment on in-memory computing (IMC) accelerators is constrained by the repeated read-modify-write (RMW) operations arising from explicit weight perturbation and the prohibitive hardware footprint of random number generators (RNGs) for statistically independent per-weight perturbations. To address these challenges, we propose an implicit-perturbation ZO (IPZO) architecture in which perturbation sums computed by an event-triggered perturbation generation unit (PGU) are combined with the weighted sums produced by the IMC array, eliminating perturbation-induced RMW operations while preserving weight-stationary execution of IMC. By exploiting spike sparsity, the PGU generates and accumulates perturbation contributions only for spike-activated weight rows, reducing the required row dimension of the RNG array. An address-driven XOR recombination scheme (PGU-XOR) is further introduced to mitigate the spatial correlations caused by direct RNG reuse (PGU-Reuse). The results show that (1) PGU-XOR matches software RNGs in accuracy on Spikingformer/CIFAR-10 (76.41% vs. 76.53%) and perplexity (PPL) on SpikeGPT/WikiText-2 (54.20 vs. 53.23), whereas PGU-Reuse degrades accuracy by 9.56 percentage points and increases PPL by 11.8; (2) implemented in a TSMC 16-nm CMOS technology, PGU-XOR incurs 40.3%-46.0% area and 15.2%-48.9% energy overhead per matrix-vector multiplication relative to PGU-Reuse, yet its faster convergence reduces the total perturbation energy to 0.51x that of PGU-Reuse at iso-accuracy; (3) IPZO reduces the perturbation energy to 0.46x-0.83x that of conventional explicit weight perturbation for a batch size of B=64 and T=4 time steps, with the advantage growing as BT decreases.