发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SEED通过将解码器Transformer重解释为隐式编码器-解码器,复用验证时的上下文表示来廉价生成高质量草稿,实现高达2.7倍加速,并优于现有自推测基线。
AI 中文摘要
自推测解码通过从目标模型自身生成草稿令牌来加速大型语言模型(LLM)推理,但在草稿质量与成本之间面临尖锐的权衡。早停方法通过在中层终止计算来廉价地生成草稿,但放弃了后续层提供的更深层表示,因此草稿质量受损。多令牌预测通过从模型的最终隐藏状态输出草稿来保持草稿质量,但在每次草稿步骤中都需要为生成这些状态而付出完整前向传播的代价。我们提出自推测编码器-解码器(SEED),一种自推测方法,通过重用验证过程中已计算的深层上下文表示,廉价地获得高质量草稿。我们将标准的仅解码器Transformer重新解释为隐式编码器-解码器:前几层(编码器)构建深层上下文表示,最后几层(解码器)从中输出令牌。编码和验证合并为单个步骤:验证由完整的编码器-解码器执行,已验证前缀的上下文表示被缓存以供草稿阶段重用。因此,草稿生成非常快速:在两次验证之间,轻量级解码器自回归地生成多个令牌,每个令牌都基于缓存的表示和之前的草稿。跨多个基准的实验表明,SEED在4B规模模型上实现了高达2.7倍的平均加速,优于早停和MTP风格的自推测基线,并且比最先进的EAGLE-3快28%,同时保持甚至提高了标准自回归微调的生成质量。代码可在该https URL获取。
英文摘要
Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-decoder (SEED), a self-speculative method that obtains high-quality drafts cheaply by reusing the deep contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder-decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder-decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the lightweight decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.7$\times$ average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 28% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning. Code is available at https://github.com/lhk2004/SEED.
CommentsAccepted to NeurIPS 2026