arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30427cs.CLcs.LG

截断上限的接受直方图揭示块扩散投机解码中的搁浅加速

Ceiling-Clipped Acceptance Histograms Indicate Stranded Speed-up in Block-Diffusion Speculative Decoding

Ephrem Wu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现块扩散投机解码存在搁浅加速问题,提出 DBloom 方法对草稿器后训练扩展块大小,在多个模型和基准上提升了提交长度,且优于 JetSpec。

中文摘要 AI 辅助

投机解码借助高效的草稿模型(草稿器)提出 token,供目标模型一次性验证,在保留目标模型输出分布的前提下加快生成速度。DFlash 和 DFlare 等高接受度块扩散草稿器可在一次并行传递中填充整个块。在许多循环中,目标模型会接受整个块,因此草稿器会在验证失败前耗尽其训练的块长度范围,我们将这种未实现的接受称为搁浅加速。按每个提示或每个循环计算的平均提交长度会掩盖该现象,而接受直方图会将其表现为上限区间的峰值,即接受整个块的循环比例。我们建议在消耗训练算力前将该直方图作为预检查。在推理时单纯扩大块大小无法恢复该加速,因为一旦块超出其训练大小,草稿器的双向注意力会改变其分布,即使在早期位置也会削弱块前端的验证效果。相反,我们采用短课程在更长的块上对草稿器进行后训练,该方法称为 DBloom。在 Qwen3-8B 和 Qwen3-4B 目标模型上,将预训练的 DFlash 和 DFlare 草稿器从块大小 16 扩展到 24,在高上限基准上使每个提示的提交长度中位数提升了 +0.8 token(最高达 +1.1);若在扩展前加入连续微调,增幅可达 1.37 token。对于另一模型系列 Gemma-4-12B-IT,相同的扩展使其在全部 7 个基准上的提交长度中位数提升了 +0.41 token(Arm A),而完整的“连续微调再扩展”流程(Arm B)相比块大小 16 的草稿器,额外提升了 +0.29 至 +0.98 token。在与当代基于树的草稿器 JetSpec 的提示匹配对比中,在树预算不超过 64 节点时,DBloom 在所有基准上都能提交更多 token。

英文摘要

Speculative decoding speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, preserving the target's output distribution. High-acceptance block-diffusion drafters such as DFlash and DFlare fill an entire block in one parallel pass. In many cycles, the target accepts the whole block, so the drafter exhausts its trained block horizon before verification fails. We call this unrealized acceptance stranded speed-up. A mean committed length, per prompt or per cycle, hides it, whereas the acceptance histogram exposes it as a spike in the ceiling bin, the fraction of cycles that accept the entire block. We recommend the histogram as a preflight check before spending training compute. Naively widening the block at inference does not recover the speed-up, because once the block outgrows its training size, the drafter's bidirectional attention shifts its distribution even at early positions and erodes front-of-block verification. Instead, we post-train the drafter on a longer block with a short curriculum that emphasizes the newly exposed positions, a method we call DBloom. Expanding the pretrained DFlash and DFlare drafters from block size 16 to 24 across Qwen3-8B and Qwen3-4B targets raises the per-prompt committed length on the high-ceiling benchmarks by a median of +0.8 tokens (up to +1.1). Once continuation fine-tuning precedes expansion, the increase reaches 1.37 tokens. The same expansion also lifts committed length on all seven benchmarks for Gemma-4-12B-IT, a different model family, by a median of +0.41 tokens (Arm A), and the full continuation-then-expand pipeline (Arm B) adds +0.29 to +0.98 tokens over the same B16 drafter. In a prompt-matched comparison against JetSpec, a contemporary tree-based drafter not used in our design, DBloom commits more tokens on every benchmark at tree budgets up to 64 nodes.

发表机构

  • Advanced Micro Devices, Inc.(超威半导体公司)

机构由 AI 辅助整理,请以论文原文为准。

↑