arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从逐位置置信度到前缀调度:推测解码中的验证器跳过

From Positionwise Confidence to Prefix Scheduling: Verifier Skipping in Speculative Decoding

Haoxuan Luo, Jameson Sandler, Ferdinando Fioretto

arXiv 2608.14787首次发表:更新:

发表机构

University of Virginia(弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对推测扩散解码的验证器瓶颈问题,提出验证器跳过策略,比较不同置信信号的调度效果,在HumanEval数据集上实现了9.6%-13.5%的验证器调用节省。

AI 中文摘要

推测解码是一种通过小型草稿模型提出多个token、再由更大的目标模型并行验证,以降低自回归生成成本的主流技术。推测扩散解码(SDD)进一步通过离散扩散模型并行生成草稿块的每个位置,消除了顺序草稿过程。然而,SDD仍会对每个块调用目标模型,使得验证成为潜在瓶颈。本文指出这创造了一个新的控制维度:是否完全调用验证器。因此,我们研究验证器跳过这一有损策略,该策略会直接提交选定的草稿前缀,并探究应使用何种置信信号来调度它。有趣的是,我们的研究发现,更好的token预测器未必能产生更好的调度器:跳过需要连续的高置信前缀,而短跳过可能会导致额外的草稿轮次。为研究这种不匹配,我们在相同策略下比较了原始置信度与学习到的边际和条件生存分数,并以严格SDD、宽容策略和top-k接受作为基线。在使用DiffuCoder-7B-Instruct和Qwen3-32B的HumanEval上,所有三种置信信号在与严格SDD相同的观测pass@1下,节省了9.6%至13.5%的验证器调用。令人惊讶的是,原始置信度的节省最多;边际生存分数在大多数位置的逐位置AUROC高于原始置信度,但两种学习到的信号均未在在线场景中占据主导。我们的分析表明,验证器跳过是一个有用的新有损维度,且令人惊讶的是,其关键挑战在于前缀调度而非仅token预测。

英文摘要

Speculative decoding is a leading technique to reduce the cost of autoregressive generation by using a small drafter to propose several tokens, which are then verified in parallel by a larger target model. Speculative diffusion decoding (SDD) further removes sequential drafting by generating every position in a draft block in parallel with a discrete diffusion model. However, SDD still invokes the target on every block, leaving verification as a potential bottleneck. This paper recognizes that this creates a new control handle: whether to invoke the verifier at all. Thus, we study verifier skipping, a lossy policy that commits a selected draft prefix directly, and ask which confidence signal should schedule it. Interestingly, our study finds that better token predictors need not yield better schedulers: skips require contiguous high-confidence prefixes, while short skips can induce additional drafting rounds. To study this mismatch, we compare raw confidence with learned marginal and conditional survival scores under the same policy, using Strict SDD, lenience, and top-$k$ acceptance as baselines. On HumanEval with DiffuCoder-7B-Instruct and Qwen3-32B, all three confidence signals save $9.6\%$ to $13.5\%$ of verifier calls at the same observed pass@1 as Strict SDD. Surprisingly, raw confidence saves the most; marginal survival has higher positionwise AUROC than raw confidence at most positions, yet neither learned signal dominates online. Our analysis shows that verifier skipping is a useful new lossy axis and, surprisingly, its key challenge is prefix scheduling rather than token prediction alone.

CommentsAccepted at UncertaiNLP 2026 (non-archival). 14 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑