arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03447cs.LGcs.AI

近似投机解码

Approximate Speculative Decoding

Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无需训练的Approximate Speculative Decoding(ASD),通过带预算的最长前缀选择优化投机解码,提升了生成吞吐量与验证器接受率,且无需新草稿模型或微调。

中文摘要 AI 辅助

投机解码通过并行用目标模型验证草稿块来加速自回归生成。在标准贪心验证下,解码会在第一个与目标argmax不同的草稿token处停止,丢弃剩余的经目标评分的后缀。尽管接受此类不匹配会改变解码轨迹,但当后缀的token在已实现的前缀下仍为目标贪心时,可使其连续后缀可重用。本文提出Approximate Speculative Decoding(ASD,近似投机解码),一种无需训练的验证器,它用带预算的最长前缀选择替代二元首不匹配截断。ASD在局部目标logit遗憾门、每块异常上限及持久请求级遗憾预算的约束下接受选定的不匹配,随后重用连续的目标贪心后缀,无需额外近似决策或目标模型前向传播。ASD既不需要新的草稿模型,也无需微调,且当预算为零时,完全退化为标准贪心验证。实验表明,ASD在匹配的严格验证下,将固定工作量吞吐量提升了3.05%至15.26%,在7项Qwen3-14B + DSpark-14B任务中平均提升7.78%;在FP4转FP8兼容设置下,搭配DeepSeek-V4-Flash(284B)与DSpark时,其在GSM8K和MATH-500上的验证器端接受率也提升了约10%至16%。源代码可公开获取:this https URL

英文摘要

Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the remaining target-scored suffix. Although accepting such a mismatch changes the decoding trajectory, it can make a contiguous suffix reusable when its tokens remain target-greedy under the realized prefix. In this paper, we introduce \textbf{Approximate Speculative Decoding (ASD)}, a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection. ASD accepts selected mismatches subject to a local target-logit regret gate, a per-block exception cap, and a persistent request-level regret budget, then reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes. ASD requires neither a new draft model nor fine-tuning, and exactly reduces to standard greedy verification when the budget is zero. Experiments show that ASD improves fixed-workload throughput by $3.05\%$--$15.26\%$ over matched strict verification and averages a $7.78\%$ gain across seven Qwen3-14B + DSpark-14B tasks. On DeepSeek-V4-Flash (284B) with DSpark it also raises verifier-side acceptance by roughly $10\%$--$16\%$ on GSM8K and MATH-500 in an FP4-to-FP8 compatibility setting. The source code is publicly available at: https://github.com/Kissmetothemoon/ASD

↑