发表机构
National University of Singapore; Shanghai Advanced Research Institute, Chinese Academy of Science; University of Chinese Academy of Sciences; Shandong University(新加坡国立大学; 中国科学院上海高等研究院; 中国科学院大学; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Falcon框架,通过流水线延迟搜索与量化感知训练,在模拟存内计算下优化脉冲神经网络延迟,证明局部等待增加可能同时降低速度与精度,并在GSCV2和SSC上以低延迟实现高精度。
AI 中文摘要
脉冲神经网络因其低功耗特性在语音命令识别领域颇具吸引力,然而其延迟问题所受到的关注远不及能效,且其多时间步执行方式被广泛认为使其比量化神经网络更慢。本文挑战了“更多局部时间步必然意味着更高网络延迟”这一假设。通过在时间步级别上重叠相邻层的计算,脉冲神经网络(SNN)可在比同类位串行量化神经网络(QNN)更短的时间内完成执行。然而,这种重叠依赖于在不完整输入上发放的脉冲,而脉冲一旦产生便无法撤回,因此其误差会持续存在并降低精度。在发放前等待更多输入似乎会以牺牲重叠度为代价来提高精度。然而,我们发现并证明这一直觉在某些层上并不成立,在这些层中,即使等待时间的小幅增加也可能改变脉冲时序和下游计算,从而使网络既更慢又更不精确。因此,我们提出了一种流水线延迟搜索方法,通过平衡任务级精度提升与新增网络延迟来为每一层选择延迟。随后,我们通过基于脉冲的量化感知训练以及对发放阈值和初始膜电位的受限调优来适配所选的配置。这些步骤共同构成了Falcon,一个用于细粒度延迟分析与受控发放的框架,该系统在具有共享数字引擎的空间模拟存内计算映射下,系统性地分析和优化SNN延迟。我们在GSCV2和SSC上评估了Falcon,在建模的网络核心延迟分别为119.64微秒和124.00微秒时,取得了具有竞争力的96.31和83.02的精度。综合我们的分析和结果,表明SNN能够计算更多却完成得更快,等待更久却预测得更差,这凸显了Falcon对延迟和精度两者的重要性。
英文摘要
Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.