arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06498cs.CL

DFlow:在块扩散投机解码中实现验证器信息流

DFlow: Enabling Verifier Information Flow in Block Diffusion Speculative Decoding

  • Peking University(北京大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Institute of Computational Social Science, Peking University (Qingdao)(北京大学(青岛)计算社会科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yaojie Zhang, Linfeng Zhang, Bin Cui, Xupeng Miao

AI总结:

DFlow通过重用验证器在拒绝位置上的隐藏状态并采用自条件训练,使信息跨起草轮次流动,从而在Qwen3模型上提升块扩散投机解码的起草质量与接受长度。

AI中文摘要:

块扩散投机解码通过并行提出一批未来令牌,并通过目标模型的一次前向传递对其进行验证,从而提升大语言模型的推理效率。然而,现有方法仅保留被接受的词缀并丢弃被拒绝的后缀,这使得这些位置上的计算无法惠及后续的起草轮次,并迫使起草者反复从头为未来令牌重建表示。我们观察到,拒绝仅决定所提议的令牌是否可以被提交,而验证器在拒绝位置上的表示仍可为后续预测提供有用信息。基于这一观察,我们提出了DFlow,一个简单而有效的框架,使验证器信息能够在各起草轮次之间流动。DFlow重用目标验证器为被拒绝后缀生成的隐藏状态,以指导后续起草,而无需额外的目标计算。为了有效学习跨起草轮次的信息流,我们引入了一种自条件训练策略,将先前预测中的验证器表示反馈到后续预测中。在Qwen3模型上跨多个基准的实验表明,与DFlash相比,DFlow持续提高了起草质量和接受长度。

英文摘要:

Block diffusion speculative decoding improves LLM inference efficiency by proposing a block of future tokens in parallel and verifying them with a single forward pass through the target model. However, existing methods retain only the accepted prefix and discard the rejected suffix, preventing the computation spent on these positions from benefiting subsequent drafting rounds and forcing the drafter to repeatedly reconstruct representations for future tokens from scratch. We observe that rejection only determines whether a proposed token can be committed, while the verifier representations at rejected positions can still provide useful information for subsequent predictions. Based on this observation, we propose DFlow, a simple yet effective framework that enables verifier information to flow across drafting rounds. DFlow reuses the hidden states produced by the target verifier for the rejected suffix to guide subsequent drafting without additional target computation. To effectively learn this information flow across drafting rounds, we introduce a self-condition train strategy that feeds verifier representations from earlier predictions back into subsequent predictions. Experiments on Qwen3 models across diverse benchmarks demonstrate that DFlow consistently improves draft quality and acceptance length over DFlash.

↑