发表机构
King’s College London; Institute for Intelligent Networked Systems (INSI), Northeastern University London(伦敦国王学院; 伦敦东北大学智能网络系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对自回归大语言模型推理低效问题,提出神经形态掩码扩散语言模型,结合块扩散与脉冲神经形态计算,利用脉冲诱导稀疏性减少参数流量和计算,开发模型分析协同效应,实验表明该模型在能源效率和吞吐量上有显著提升。
AI 中文摘要
自回归(AR)大语言模型在推理时本质上效率低下,因为每个生成的令牌都需要访问完整的模型参数集,导致低运算强度和高能耗。掩码扩散语言模型(MDLM)通过允许每次参数访问生成多个令牌,部分解决了内存受限设置的这一限制。为了进一步提高在具有大量片上内存的现代平台上的推理效率,本文提出了神经形态MDLM(N-MDLM),它将块扩散与基于脉冲的神经形态计算相结合,共同提高吞吐量和能源效率。块扩散通过每次参数访问生成多个令牌来提高令牌吞吐量,而脉冲诱导的稀疏性通过跳过非活动通道来减少有效参数流量和计算量。为了分析稀疏性和扩散的协同效应,我们开发了一个受令牌级屋顶线启发的模型,该模型捕捉了块并行生成和脉冲稀疏性对解码效率的综合影响。在翻译任务上的实验结果表明,由于脉冲诱导的稀疏性,即使在计算受限的平台上,N-MDLM在能源效率和吞吐量方面也取得了显著提高,而MDLM在这些平台上无法比AR-LLM有更好的提升。
英文摘要
Autoregressive (AR) large language models (LLMs) are inherently inefficient at inference time because each generated token requires accessing the full set of model parameters, leading to low operational intensity and high energy consumption. Masked diffusion language models (MDLMs) partially address this limitation for memory-bound settings by allowing multiple tokens to be generated per parameter access. In order to further enhance inference efficiency on modern platforms with extensive in-chip memory, this work proposes neuromorphic MDLMs (N-MDLMs), which integrate block diffusion with spike-based neuromorphic computation to jointly improve throughput and energy efficiency. While block diffusion increases token throughput by producing multiple tokens per parameter access, spike-induced sparsity reduces effective parameter traffic and computations by skipping inactive channels. To analyze the synergistic effect of sparsity and diffusion, we develop a token-level roofline-inspired model that captures the combined impact of block-parallel generation and spike sparsity on decoding efficiency. Experimental results on translation tasks show that, thanks to spike-induced sparsity, N-MDLMs achieve substantial improvements in energy efficiency and throughput even in compute-bound platforms for which MDLMs would fail to improve over AR-LLMs.
CommentsAccepted for presentation at 2026 IEEE Workshop on Signal Processing Systems (SiPS)