arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

D-Loop:用于投机解码的循环扩散草稿生成

D-Loop: Looped Diffusion Drafting for Speculative Decoding

Kecheng Chen, Yuyang He, Cheng Gong, Hui Liu, Guoping Long, Jiajun Li, Shi Wu, Suiyun Zhang, Haoliang Li, Ziru Liu, Rui Liu

arXiv 2610.06011首次发表:更新:

发表机构

City University of Hong Kong; The Chinese University of Hong Kong; Huawei Research(香港城市大学; 香港中文大学; 华为研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

D-Loop通过块内因果条件化和循环参数共享,在不增加模型组件的情况下改进扩散草稿生成,解决重复陷阱,提升投机解码的接受长度,并在多个基准上超越现有方法。

AI 中文摘要

块扩散通过一次前向传播草拟多个标记来加速投机解码。然而,每个位置预测边际分布时并未观察先前提出的标记,这限制了草稿质量和接受长度。我们识别出一个具体失败模式——“重复陷阱”,即相邻位置产生冗余的相同标记副本。我们从理论上解释这一趋势,并通过实验检验其与较短接受草稿的关联。近期方法通过额外的因果头或单独训练的草稿模型来细化边际预测,但这增加了参数存储并引入了单独的训练目标。我们转而提出D-Loop,它在原始扩散草稿模型内部引入“块内因果条件化”,无需额外模型组件。受半自回归生成和参数共享的启发,D-Loop在循环传递中复用同一骨干网络。第一次传递提出一个块,第二次传递基于选定前缀并行重新生成后缀。一个互补的前缀-后缀目标训练共享草稿模型,使其既能进行仅锚点的前缀预测,也能进行前缀条件化的后缀预测。在八个数学、代码和聊天基准测试中,D-Loop在Qwen3-4B和Qwen3-8B上均能超越DFlash和DSpark,并取得显著提升。

英文摘要

Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑