发表机构
Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对传统TTFS SNN难以编码LLM部分模块的问题,提出基于参考的策略构建全TTFS SNN架构,在BERT、GPT-2上验证其性能,首次将TTFS编码的脉冲LLM扩展至15亿参数。
AI 中文摘要
脉冲神经网络(SNN)凭借其固有的稀疏事件驱动计算,为构建高能效大语言模型(LLM)提供了可行路径。首次脉冲时间(TTFS)编码在一个时间窗口内为每个神经元生成至多一个脉冲,从而产生极低的脉冲发放率。然而,传统的TTFS SNN受限于特定结构,难以使用TTFS编码LLM中的某些模块,如层归一化和矩阵乘法。为克服这一局限,本文提出一种基于参考的策略,专门用于编码LLM的四个核心组件:嵌入层、层归一化、注意力相关操作和Dropout。我们构建了一个完全基于TTFS的SNN架构并对其进行端到端训练。在BERT和GPT-2等现代LLM上开展的实验表明,本文方法在自然语言理解和常识推理任务上的性能可与人工神经网络(ANN) counterpart相媲美,但在语言建模困惑度方面仍存在明显差距。据我们所知,这是首个使用TTFS编码将脉冲LLM扩展至15亿参数的研究。我们还报告了与脉冲相关的能量估计值,该值是在既定成本模型下的脉冲计数代理,而非神经形态硬件上的实际测量值。
英文摘要
Leveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising path toward energy-efficient large language models (LLMs). Time-to-first-spike (TTFS) coding generates at most one spike per neuron within a time window, yielding extremely low firing rates. However, conventional TTFS SNNs are restricted to specific structures, making it challenging to encode certain blocks in LLM -- such as layer normalization and matrix multiplication --using TTFS. To overcome this limitation, we introduce a reference-based strategy specifically to encode the four core LLM components: embedding layers, layer normalization, attention-related operations and dropout. We construct a fully TTFS-based SNN architecture and train it end-to-end. Experiments on modern LLMs like BERT and GPT-2 demonstrate that our approach achieves performance comparable to ANN counterparts on natural language understanding and common-sense reasoning, while a clear gap remains on language modeling perplexity. To the best of our knowledge, this is the first work to scale a spiking LLM to 1.5 billion parameters using TTFS coding. We also report an estimate of spike-related energy; this is a spike-count proxy under an established cost model rather than a measurement on neuromorphic hardware.