并行时间对齐尖峰自注意力:用于一致的整数值训练与尖峰驱动推理
Parallel Time-Aligned Spiking Self-Attention for Consistent Integer-Valued Training and Spike-Driven Inference
浏览论文内容
中文总结 AI 辅助
针对尖峰自注意力训练与推理不一致问题,提出并行时间对齐尖峰自注意力(PT-SSA)及自适应变体,显著降低Top-1差距并提升吞吐量。
中文摘要 AI 辅助
整数值泄漏积分激发(I-LIF)神经元和尖峰激发近似(SFA)分别通过将尖峰序列表示为激发计数和归一化激发率来降低时间训练成本。然而,将尖峰自注意力(SSA)直接应用于这些压缩的查询、键和值表示会引入在尖峰驱动推理期间不存在的跨时间交互。我们将这种算子级差异称为时间交互不匹配(TIM)。我们提出了并行时间对齐尖峰自注意力(PT-SSA),它从I-LIF计数或SFA激发率重建连续的虚拟尖峰切片,仅在时间对齐的切片之间并行计算注意力,并对每步输出求和。为了适应SFA下注意力输出规模的减小,我们进一步引入了自适应PT-SSA,它在输出SFA神经元之前学习每个块的正缩放,以提高激发级利用率。在CIFAR-10、CIFAR-100和ImageNet-1K上的实验表明,所提出的方法显著减少了训练与推理之间的不匹配。在CIFAR-100上使用I-LIF时,PT-SSA将平均Top-1差距从1.85个百分点降至0.29个百分点。在ImageNet-1K上,自适应PT-SSA将Top-1差距从27.78个百分点降至0.06个百分点,并实现了74.53%的尖峰驱动Top-1准确率。一个Triton融合的PT-SSA训练内核相对于循环LIF SSA保持了2.91倍的吞吐量优势。
英文摘要
Integer-valued leaky integrate-and-fire (I-LIF) neurons and spike firing approximation (SFA) reduce temporal training cost by representing spike trains as firing counts and normalized firing rates, respectively. However, applying spiking self-attention (SSA) directly to these compressed query, key, and value representations introduces cross-time interactions that are absent during spike-driven inference. We term this operator-level discrepancy Temporal Interaction Mismatch (TIM). We propose Parallel Time-Aligned Spiking Self-Attention (PT-SSA), which reconstructs consecutive virtual spike slices from either I-LIF counts or SFA firing rates, computes attention only between time-aligned slices in parallel, and sums the per-step outputs. To accommodate the reduced attention output scale under SFA, we further introduce Adaptive PT-SSA, which learns a positive per-block rescaling before the output SFA neuron to improve firing-level utilization. Experiments on CIFAR-10, CIFAR-100, and ImageNet-1K show that the proposed methods substantially reduce train--inference mismatch. On CIFAR-100 with I-LIF, PT-SSA reduces the mean Top-1 gap from 1.85 to 0.29 percentage points. On ImageNet-1K, Adaptive PT-SSA reduces the Top-1 gap from 27.78 to 0.06 percentage points and achieves 74.53\% spike-driven Top-1 accuracy. A Triton-fused PT-SSA training kernel retains a $2.91\times$ throughput advantage over recurrent LIF SSA.
发表机构
- Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
- Pengcheng Laboratory(鹏城实验室)
- Shenzhen Graduate School, Peking University(北京大学深圳研究生院)
- School of Computer Science, Peking University(北京大学计算机学院)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。