发表机构
Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QuantaSpike提出基于LTIF神经元的短窗口脉冲驱动量化框架,通过组自适应增益和选择性异常值接纳,在OPT和Llama-2上实现先进性能,并大幅降低推理能耗。
AI 中文摘要
大型语言模型(LLM)在众多任务上表现出色,但在推理过程中依赖密集的乘加(MAC)运算,导致高能耗。脉冲神经网络(SNN)提供了一种事件驱动的替代方案,其中突触整合使用轻量级累加。然而,脉冲驱动的LLM推理仍然困难,因为异常值密集的激活通常需要长脉冲窗口或辅助的非脉冲路径。我们提出了QuantaSpike,一个围绕对数三值积分-发放(LTIF)神经元构建的短窗口脉冲驱动量化框架。LTIF使用具有2的幂次膜响应量子的三值事件,提高了每个发放步骤所表示的信息,同时保持移位累加(shift-ACC)兼容的计算。QuantaSpike将该神经元与组自适应增益和选择性异常值接纳相结合:正常值使用残差LTIF步骤,而被接纳的异常值在进入相同的残差动态之前接收一个额外的起始脉冲。在OPT和Llama-2上,QuantaSpike在脉冲驱动LLM量化方法中实现了最先进或具有竞争力的困惑度和零样本准确率。它还能迁移到更新的密集LLM,在Llama-3-8B和Qwen3-8B上在相同的四步脉冲窗口下保持接近FP16参考。解析线性能量预测显示,相对于SpikeQuant,QuantaSpike在OPT模型上将一次线性变换的能量减少了约80.0%,在Llama-2模型上减少了67.1%,为LLM推理提供了一条准确且节能的脉冲驱动路径。
英文摘要
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about $80.0\%$ on OPT models and $67.1\%$ on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.