arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14717cs.LGcs.AI

Carryover Drafting: 回收被拒绝状态用于投机解码

Carryover Drafting: Recycling Rejected States for Speculative Decoding

Jahyun Koo, Sunghyeon Woo, Jaeeun Kil, Jeongtae Lee, Sungjae Lee, Kyomin Jung, Minsub Kim

首次发表
浏览论文内容

中文总结 AI 辅助

针对投机解码中被拒绝令牌计算浪费的问题,提出Carryover Drafting机制,回收被拒绝隐藏状态作为临时KV上下文,配合并行草稿-验证-草稿训练,提升接受长度与端到端推理速度。

中文摘要 AI 辅助

投机解码通过并行验证多个草稿令牌来加速大语言模型推理,使得单次目标前向传播能够接受多个令牌。根据构造,验证过程会同时计算被接受和被拒绝令牌的表示。然而,传统的草稿器仅保留被接受令牌的表示,导致花费在被拒绝令牌上的大量验证计算被有效浪费。我们发现,目标前向过程中产生的这些被丢弃的隐藏状态保留了关于未来令牌的有用信息,可以改进后续的草稿。然而,实现这一机会面临两个不同的挑战。在推理时,回收开销可能增加草稿延迟,从而削弱因接受长度增加而带来的加速。在训练时,标准的并行草稿器训练不会产生与推理对齐的被拒绝状态,而通过顺序展开获取这些状态则会牺牲训练位置间的并行性。我们提出了Carryover Drafting,它解决了这两个挑战。Carryover将被拒绝的目标隐藏状态回收作为临时KV上下文,使草稿器能够选择性地关注它们。它重用草稿器现有的接口,仅添加一个学习到的嵌入来区分被拒绝状态与已提交上下文。额外的KV上下文在每个草稿轮次中被替换,其长度保持在一个提案块内。我们引入了并行草稿-验证-草稿训练,使草稿器暴露于与推理对齐的被拒绝状态,同时保持训练位置间的并行性。使用DFlash和源自DSpark的半自回归草稿器在两种目标模型上的实验表明,这种简单的Carryover机制相比相应基线,平均接受长度提高了6.5%至14.7%,端到端vLLM加速提高了7.9%至14.4%,在翻译任务上加速提升达到28.8%。

英文摘要

Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to accept several tokens. By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted. We find that these discarded hidden states generated during target forward retain useful information about future tokens that can improve subsequent drafts. However, realizing this opportunity poses two distinct challenges. At inference, recycling overhead can increase drafting latency, diminishing the speedup gained from increased acceptance length. During training, standard parallel drafter training does not produce inference-aligned rejected states, while obtaining them through sequential rollouts would sacrifice parallelism across training positions. We introduce Carryover Drafting, which addresses both challenges. Carryover recycles rejected target hidden states as temporary KV context, allowing the drafter to selectively attend to them. It reuses the drafter's existing interface and adds only a single learned embedding to distinguish rejected states from committed context. The additional KV context is replaced each drafting round, keeping its length bounded by one proposal block. We introduce parallel draft--verify--draft training that exposes the drafter to inference-aligned rejected states while preserving parallelism across training positions. Experiments with DFlash and a DSpark-derived semi-autoregressive drafter across two target models show that this simple Carryover mechanism improves average acceptance length by 6.5--14.7% and end-to-end vLLM speedup by 7.9--14.4% over the corresponding baselines, with speedup gains reaching 28.8% on translation.

发表机构

  • NAVER Cloud(NAVER云)
  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑