DBLAST:面向随机投机解码的依赖块草稿生成
DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding
浏览论文内容
中文总结 AI 辅助
本文针对块扩散草稿模型在高熵投机解码下可接受草稿长度下降的问题,提出DBLast依赖块草稿模型,在多基准测试中相比独立块采样稳定提升可接受长度。
中文摘要 AI 辅助
投机解码通过使用轻量级草稿模型生成多个未来token、目标模型进行验证,来加速大语言模型的推理。近期的块式和扩散式草稿模型虽能单次预测多个位置,但其训练和采样流程通常针对贪心解码优化,或假设草稿块内的位置条件独立。该假设在非贪心投机解码中会失效,此时目标分布刻意随机,存在多个合理延续。本文研究了块扩散草稿模型的这种不匹配问题,表明当目标采样分布的熵增加时,可接受的草稿长度会下降。我们提出了一种基于token位置低秩潜在混合的依赖块草稿模型,辅以面向接受度的训练目标,直接以期望验证长度为优化目标。在GSM8K、MT-Bench、HumanEval和创意写作基准上,使用Qwen3-4B和Qwen3-8B的实验显示,我们的方法DBLast相比独立块采样,可稳定提升可接受长度,尤其在高熵解码场景中效果显著。
英文摘要
Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them. While recent block and diffusion-style drafters can predict several positions in a single pass, their training and sampling procedures are typically optimized for greedy decoding or assume that positions in the draft block are conditionally independent. This assumption becomes brittle in non-greedy speculative decoding, where the target distribution is deliberately stochastic and multiple continuations become plausible. We study this mismatch for block diffusion drafters and show that the accepted draft length degrades as the entropy of the target sampling distribution increases. We propose a dependent block drafter based on a low-rank latent mixture over token positions, complemented by an acceptance-oriented training objective that directly targets the expected verified length. Experiments with Qwen3-4B and Qwen3-8B on GSM8K, MT-Bench, HumanEval, and creative-writing benchmarks show that our approach, namely DBLast, consistently improves accepted length over independent block sampling, especially in higher-entropy decoding regimes.