DLoop: 循环推测解码
DLoop: Looped Speculative Decoding
查看机构详情
- NAVER AI Lab(NAVER人工智能实验室)
- NAVER AI Search Platform(NAVER人工智能搜索平台)
- Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
DLoop提出循环推测解码,通过多草稿阶段自适应累积标记并统一验证,减少目标模型前向传播,在多种方法上提升5%-41%速度且无损。
中文摘要 AI 辅助
推测解码加速了大型语言模型中的自回归生成。在每个草稿阶段,一个轻量级草稿模型提出标记,随后由目标模型进行验证。随着草稿模型能力的日益增强,我们发现目标模型经常接受草稿阶段产生的所有标记。然而,每个草稿阶段之后仍会进行验证,导致即使草稿可以继续,也会产生不必要的目标模型前向传播。自适应草稿长度方法在解码过程中决定验证之前有多少个草稿标记,但这种方法仅提高了自回归草稿模型的速度。对于并行草稿模型,草稿进一步需要目标模型对未验证草稿标记的隐藏状态。我们提出了DLoop,一种循环形式的推测解码,它在验证之前自适应地执行多个草稿阶段。当草稿模型保持自信时,DLoop继续草稿,并一起验证所有累积的草稿标记。循环感知训练通过将草稿模型暴露于其自身对未验证草稿标记的隐藏状态,使其在额外的草稿阶段保持可靠。通过花费额外的草稿模型前向传播,DLoop减少了验证所需的目标模型前向传播次数。在包括EAGLE-3、DFlash、Domino、DSpark和多标记预测模块在内的多种推测解码方法中,DLoop将墙钟加速提高了5%至41%,同时保持了无损解码。代码将在提供的https URL上提供。
英文摘要
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at https://github.com/naver-ai/DLoop.