arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向神经形态硬件的、具有稀疏神经活动的事件驱动语言模型

Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

Simon Richter, Ruhai Lin, Jason Yik, Taylor Kergan, Rui-Jie Zhu, Farshad Moradi, Jason Eshraghian

arXiv 2608.30439首次发表:更新:

发表机构

Aarhus University; University of California, Santa Cruz; Harvard University; University of Southern Denmark(奥胡斯大学; 加利福尼亚大学圣克鲁兹分校; 哈佛大学; 南丹麦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对Transformer大语言模型推理的内存与计算瓶颈,提出在重度量化线性注意力模型中引入稀疏神经活动的方法,可提升神经形态平台的LLM部署性能,实现吞吐量与功耗的显著优化。

AI 中文摘要

基于Transformer的大语言模型(LLM)的推理常受限于内存受限的KV缓存和二次方注意力成本。状态空间模型(SSM)通过线性注意力和固定大小的循环状态缓解了这一问题,但即使经过量化,其大型稠密线性投影的计算成本仍然很高。我们提出一种方法,在重度量化的线性注意力模型中诱导稀疏神经活动,同时保持极小的性能损失。低于每个投影可训练阈值(±Δ)的激活会被置零,同时保留关键的异常值,实现与稠密模型相当的性能,有效算术运算减少多达4倍。针对多核、多芯片神经形态平台,事件驱动执行在计算和通信层面将非结构化稀疏性转化为吞吐量,这是GPU架构根本不具备的能力;我们预计,与基于Transformer的可比模型的边缘GPU推理相比,吞吐量可提升多达37倍,功耗可降低16倍,与非稀疏化基线相比可提升5.4倍。这些结果表明,稀疏、量化的线性注意力模型非常适合在事件驱动的多核平台上部署LLM。

英文摘要

Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space models (SSMs) mitigate this through linear attention and fixed-size recurrent states, but their large dense linear projections remain computationally expensive even after quantization. We introduce a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss. Activations below a per-projection trainable threshold ($\pm Δ$) are nullified while preserving crucial outliers, achieving comparable performance to dense models with up to 4$\times$ fewer effective arithmetic operations. Targeting a multi-core, multi-chip neuromorphic platform, where event-driven execution converts unstructured sparsity into throughput at both the compute and communication levels, a capability GPU architectures fundamentally lack, we project up to 37$\times$ higher throughput and 16$\times$ lower power versus edge GPU inference of a comparable transformer-based model, and up to 5.4$\times$ improvements over the non-sparsified baseline. These results position sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.

Comments8 pages, 4 figures, Accepted at IEEE MCSOC2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑