arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38510cs.CL

DEdit:用于投机解码的迭代草稿编辑

DEdit: Iterative Draft Editing for Speculative Decoding

Longxuan Yu, Bingsen Chen, Peng Shi, Dongkyu Lee, Yi Xiang, Hideo Kobayashi, Sheng Zhang, Shuaichen Chang, Xing Niu, Zhuoyan Xu, Greg Ver Steeg, Jiarong Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

DEdit是一种基于扩散的草稿模型,通过迭代编辑草稿修复早期错误,结合ProposalMix训练,在Qwen3模型上实现最高加速比,显著提升投机解码效率。

中文摘要 AI 辅助

投机解码通过让轻量级草稿模型提出令牌,并由目标模型并行验证,从而加速自回归大语言模型。基于扩散的草稿模型通过一次提出多个令牌进一步减少草稿延迟。然而,这些令牌是独立预测的,因此单个早期错误会导致前缀验证丢弃草稿的其余部分,即使其中包含有用的下游预测。我们引入了DEdit,一种基于扩散的草稿模型,它不仅可以通过传统的并行去掩蔽进行草稿,还可以通过令牌到令牌的预测迭代地编辑其草稿。通过编辑,后续预测可以作为双向上下文,用于修复早期错误并扩展接受的前缀。为了教会模型在保留正确预测的同时修复错误,我们提出了ProposalMix,一种训练方案,在训练期间根据首轮置信度将草稿预测与真实令牌混合。在Qwen3-4B和Qwen3-8B上的七个基准测试中,DEdit在评估的草稿模型中实现了最高的宏平均令牌接受率和加速比,在贪婪解码下分别达到了相对于自回归生成的宏平均加速比5.72倍和5.97倍。进一步分析表明,随着更多编辑轮次和更宽的草稿窗口,接受率提高,并且ProposalMix将缩短接受前缀的有害编辑减半。此外,将编辑器限制为因果注意力会降低接受率,尤其是在高度可预测的输出上,这表明未来上下文是这些收益的关键来源。

英文摘要

Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions. We introduce DEdit, a diffusion-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token-to-token predictions. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix. To teach the model to repair errors while preserving correct predictions, we propose ProposalMix, a training scheme that mixes draft predictions with ground-truth tokens based on first-pass confidence during training. Across seven benchmarks on Qwen3-4B and Qwen3-8B, DEdit achieves the highest macro-average token acceptance and speedup among the evaluated drafters, reaching macro-average speedups of $5.72\times$ and $5.97\times$ over autoregressive generation under greedy decoding, respectively. Further analysis shows that acceptance improves with more editing passes and wider drafting windows, and that ProposalMix halves harmful edits that shorten the accepted prefix. Moreover, restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating that future context is a key source of these gains.

发表机构

  • University of California, Riverside(加州大学河滨分校)
  • Amazon Web Services(亚马逊云服务)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑