发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对线性注意力模型自回归解码效率低的问题,提出SpecLA推测解码运行时,通过拓扑感知内核验证、存储紧凑因子及置信度修剪等方法,在NVIDIA H100上实现了比自回归解码高达1.70倍的端到端加速。
AI 中文摘要
线性注意力模型用循环状态取代不断增长的KV缓存,但自回归解码仍一次一个令牌地读取、更新和写入这些状态。推测解码可通过在一次目标传递中验证多个草稿令牌来降低成本,但现有推测系统是为Transformer KV缓存设计的。对于有状态线性注意力目标,验证必须遵循跨链和分支的循环依赖,接受必须仅更新接受的状态轨迹,起草者必须避免提交浪费有状态验证工作的候选。本文提出SpecLA,一种有状态线性注意力模型的推测解码运行时。SpecLA用拓扑感知内核验证链和树,存储验证期间产生的紧凑因子以恢复接受状态,并使用置信度修剪和目标对齐的EAGLE风格起草者向验证器提供有用候选。在具有公共GDN-1.3B目标的NVIDIA H100上,SpecLA比自回归解码实现高达1.70倍的端到端加速。
英文摘要
Linear-attention models replace the growing KV cache with recurrent states, but autoregressive decoding still reads, updates, and writes these states one token at a time. Speculative decoding can reduce this cost by verifying several draft tokens in one target pass, yet existing speculative systems are designed for Transformer KV caches. For stateful linear-attention targets, verification must follow recurrent dependencies across chains and branches, acceptance must update only the accepted state trajectory, and the drafter must avoid submitting candidates that waste stateful verification work. This paper presents SpecLA, a speculative decoding runtime for stateful linear-attention models. SpecLA verifies chains and trees with topology-aware kernels, stores compact factors produced during verification to recover accepted states, and uses confidence pruning plus a target-aligned EAGLE-style drafter to feed useful candidates to the verifier. On an NVIDIA H100 with a public GDN-1.3B target, SpecLA achieves up to 1.70x end-to-end speedup over autoregressive decoding.