arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.22634cs.AI

PRESTO:用于扩散推测解码的前缀对齐树草稿生成

PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Georgia Institute of Technology(佐治亚理工学院)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Zheng Wang, Zhifan Ye, Qi Cheng, Yonggan Fu, Ziyan Wang, Feng Zhu, Haozhe Zhao, Jan Kautz, Pavlo Molchanov, Humphrey Shi, Minjia Zhang

AI总结:

研究针对扩散大语言模型用于推测解码时现有方法的局限,提出PRESTO框架,通过前缀对齐评分和基于优先级的树搜索,解决扩散草稿与前缀验证不匹配问题,在多种基准测试中显著提升了端到端吞吐量。

AI中文摘要:

扩散大语言模型(dLLMs)已成为自回归(AR)大语言模型的一个有前途的替代方案,能够并行生成令牌。这使其成为推测解码(SD)的有效草稿模型,可在一次前向传播中生成整个草稿令牌块。然而,现有的基于扩散的草稿生成方法依赖线性草稿生成,尽管dLLMs在各个位置会发出多个候选令牌,这导致了一个庞大的解码路径组合空间,限制了接受长度和解码效率。为利用这种多候选结构,我们将基于树的草稿生成应用于扩散草稿生成器,以探索不同的候选路径。但我们发现简单的树草稿生成并非最优:扩散边缘概率是前缀盲的,与基于前缀的AR验证不匹配,导致路径排名不可靠。我们提出了PRESTO,这是一个有原则的框架,通过前缀对齐评分和基于优先级的树搜索来解决扩散草稿置信度与基于前缀的AR验证之间的根本不匹配问题,将基于树的草稿生成扩展到扩散草稿生成器。PRESTO背后的关键原则是:(1)候选排名应与基于前缀的AR验证性质对齐;(2)树构建应优先考虑具有高验证潜力的候选路径,以最大化接受长度。大量实验表明,在各种基准测试中,PRESTO在最先进的专用扩散草稿生成器SD上实现了平均高达1.5倍的端到端吞吐量加速,在自推测扩散大语言模型上实现了平均1.12倍的加速。

英文摘要:

Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel. This makes them effective draft models for speculative decoding (SD), producing an entire block of draft tokens in a single forward pass. Yet existing diffusion-based drafting methods rely on linear drafting, even though dLLMs emit multiple candidate tokens across positions, inducing a large combinatorial space of decoding paths. Consequently, they limit acceptance length and decoding efficiency. To exploit this multi-candidate structure, we apply tree-based drafting to diffusion drafters, enabling exploration of diverse candidate paths. However, we find that naive tree drafting is suboptimal: diffusion marginals are prefix-blind, mismatching the prefix-based AR verification and yielding unreliable path ranking. We propose PRESTO, a principled framework that extends tree-based drafting to diffusion drafters while resolving the fundamental mismatch between diffusion draft confidence and prefix-based AR verification through PREfix-aligned Scoring and priority-based Tree search for diffusion speculative decOding. The key principles behind PRESTO are that (1) candidate ranking should align with the prefix-based nature of AR verification, and (2) tree construction should prioritize candidate paths with high verification potential to maximize acceptance length. Extensive experiments show that PRESTO achieves up to an average of $1.5\times$ end-to-end throughput speedup on the state-of-the-art dedicated diffusion drafter SD and an average of $1.12\times$ on self-speculative diffusion LLMs across diverse benchmarks.

↑