arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于可微脉冲时间离散化的连续时间脉冲神经网络的可扩展训练

Scalable Training of Continuous-Time Spiking Neural Networks with Differentiable Spike-Time Discretization

Yusuke Sakemi, Tomoya Takeuchi, Takeo Hosomi, Kazuyuki Aihara

arXiv 2607.14672首次发表:更新:

发表机构

Research Center for Mathematical Engineering, Chiba Institute of Technology; International Research Center for Neurointelligence (WPI-IRCN), The University of Tokyo; NEC Corporation(千叶工业大学数学工程研究中心; 东京大学国际神经智能研究中心(WPI-IRCN); 日本电气公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究连续时间脉冲神经网络训练受内存限制问题,提出基于可微脉冲时间离散化的内存高效训练框架及受同步发放链启发的时间正则化,降低内存消耗与训练时间,实现多层卷积SNNs在单GPU上训练。

AI 中文摘要

连续时间脉冲神经网络(SNN)为时间计算、计算神经科学和神经形态硬件提供了一个事件驱动框架。然而,训练深度连续时间SNNs受到精确脉冲时间计算所需内存的严重限制。本文引入了一种基于可微脉冲时间离散化(DSTD)的内存高效训练框架,用于具有一般膜和突触时间常数的泄漏积分发放神经元。DSTD将不规则的突触前脉冲映射到固定时间点的可微加权事件上,在精确逼近连续时间膜电位动力学的同时,将与候选相关的激活内存从$O(N_{\mathrm{out}}N_{\mathrm{in}})$减少到$O(N_{\mathrm{out}}M)$。还引入了受同步发放链启发的时间正则化,减轻了死神经元故障,并实现了类似流水线的处理。在密集LIF层中,与精确脉冲时间计算相比,DSTD将峰值内存消耗降低了约100倍,训练时间降低了约20倍。这些方法使我们能够在单个GPU上训练CIFAR-10上的9层卷积SNNs和Fashion-MNIST上的20层卷积SNNs。

英文摘要

Continuous-time spiking neural networks (SNNs) provide an event-driven framework for temporal computation, computational neuroscience, and neuromorphic hardware. However, training deep continuous-time SNNs is severely constrained by the memory required for exact spike-time computation, which evaluates and retains candidate firing times over intervals determined by presynaptic spike ordering. Here we introduce a memory-efficient training framework based on differentiable spike-time discretization (DSTD) for leaky integrate-and-fire neurons with general membrane and synaptic time constants. DSTD maps irregular presynaptic spikes onto differentiable weighted events at fixed time points, replacing the input-dependent candidate dimension with $M$ fixed time intervals while accurately approximating continuous-time membrane-potential dynamics. This reduces candidate-related activation memory from $O(N_{\mathrm{out}}N_{\mathrm{in}})$ to $O(N_{\mathrm{out}}M)$ in the case of time-to-first-spike (TTFS) coding, where $N_{\mathrm{in}}$ and $N_{\mathrm{out}}$ denote the numbers of presynaptic and postsynaptic neurons, respectively. We further introduce synfire-chain-inspired temporal regularization that organizes layer-wise firing windows, mitigates dead-neuron failures, and enables pipeline-like processing. In dense LIF layers, DSTD reduced peak memory consumption by up to approximately 100-fold and training time by up to approximately 20-fold compared with exact spike-time computation. Together, these methods allowed us to train 9-layer convolutional SNNs on CIFAR-10 and 20-layer convolutional SNNs on Fashion-MNIST on a single GPU.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑