arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13524cs.LG

DARTree:基于自回归草稿树的推测式扩散解码

DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

Tianyi Li, Yaxin Luo, Xinyi Shang, Zhiqiang Shen

首次发表
浏览论文内容

中文总结 AI 辅助

DARTree是无需训练的推测式解码方法,将AR修正头从链扩展至树,在7个数学、代码、聊天基准的4种模型-温度配置下,实现最高平均接受长度与加速比,达9.73倍无损加速。

中文摘要 AI 辅助

推测式解码通过并行验证多个草稿词元,无损加速自回归语言模型。基于扩散的草稿模型进一步通过并行预测整段词元块降低提议延迟,但其位置分布是边缘分布,而非基于每条草稿路径选中的词元条件化。现有循环修正沿单条草稿链纳入因果信息,而基于扩散的树构建扩大了候选覆盖范围,却未沿各分支传递该修正。本文提出DARTree,一种无需训练的推测式解码方法,将预训练的自回归(AR)修正头从链扩展至树。DARTree首先在单批次中通过扩展和评分各深度的所有节点,构建固定宽度的候选树,随后仅应用最佳优先剪枝选择验证树,将AR头推理与顺序堆操作解耦。在7个数学、代码和聊天基准测试中,DARTree在全部4种模型-温度配置下均达到最高平均接受长度和加速比,在相同设置下每轮验证最多接受12.97个词元,比DFlash多98.6%,比Domino多27.9%,且达到比本地测量的自回归解码最高9.73倍的无损加速。

英文摘要

Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path. Existing recurrent correction incorporates causal information along a single draft chain, whereas diffusion-based tree construction broadens candidate coverage without carrying this correction along individual branches. We introduce DARTree, a training-free speculative decoding method that extends a pretrained AR correction head from chains to trees. DARTree first constructs a fixed-width candidate tree by expanding and scoring all nodes at each depth in a single batch, and then only applies best-first pruning to select the verification tree, decoupling AR-head inference from sequential heap operations. Across seven math, code, and chat benchmarks, DARTree achieves the highest average acceptance length and speedup in all four model--temperature configurations, accepting up to 12.97 tokens per verification round, 98.6\% more than DFlash and 27.9\% more than Domino in the same setting, and reaching up to 9.73$\times$ lossless speedup over locally measured autoregressive decoding.

↑