arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27735cs.CL

对整体的稀疏一瞥:无训练的自推测解码

A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding

Yuesong Liu, Yuan Zeng, Min Lyu, Ruilin Liu, Yu Guo, Yinlong Xu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对推测解码加速受限问题,提出无训练的SparseSpec-L框架,通过动态稀疏KV缓存与熵控推测长度,在保留目标模型输出分布的同时实现长上下文推理的端到端加速。

中文摘要 AI 辅助

推测解码缓解了大语言模型推理中的内存带宽瓶颈,但其加速效果同时受草稿生成开销、令牌接受率和推测长度的约束。我们提出了一项统一效率分析,结果显示当边际接受概率低于相对草稿生成成本时,延长推测范围会降低而非提升加速比。基于该分析,我们引入SparseSpec-L,这是一种用于长上下文推理的无训练自推测解码框架。SparseSpec-L利用动态稀疏化且可召回的KV缓存直接从目标模型生成轻量草稿,它将全上下文验证期间产生的每头注意力统计数据作为无额外前向传播的重要性信号,使关键历史令牌可被召回,同时无需永久丢弃密集KV缓存。在线基于熵的控制器会根据预期的逐步效率进一步选择推测长度。在多个长上下文任务和模型规模上开展的实验显示,该方法实现了一致的端到端加速,在保留目标模型输出分布的同时,达到了自回归解码的加速效果。

英文摘要

Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafting overhead, token acceptance, and speculation length. We present a unified efficiency analysis showing that extending the speculation horizon can reduce rather than improve speedup when the marginal acceptance probability falls below the relative drafting cost. Guided by this analysis, we introduce SparseSpec-L, a training-free self-speculative decoding framework for long-context inference. SparseSpec-L generates lightweight drafts directly from the target model using a dynamically sparsified and recallable KV cache. It recycles per-head attention statistics produced during full-context verification as a no-extra-forward importance signal, allowing critical historical tokens to be recalled without permanently discarding the dense KV cache. An online entropy-based controller further selects the speculation length according to expected step-wise efficiency. Experiments across multiple long-context tasks and model scales show consistent end-to-end acceleration, with up to speedup over autoregressive decoding while preserving the target model's output distribution.

补充信息

↑