发表机构
a2sys; Seoul National University; University of Wisconsin(a2sys; 首尔大学; 威斯康星大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散草稿模型训练目标与序列级验证不匹配的问题,提出块验证感知损失(BV损失),直接基于块验证接受规则最大化期望接受长度,在数学、代码和聊天基准上较交叉熵损失提升13%-21%的接受标记数。
AI 中文摘要
扩散草稿模型通过并行提出多个标记来加速推测解码。尽管近年来通过序列级草稿和验证在推测解码方面取得了进展,但现有的训练目标大多仍围绕标记级验证设计。为解决这一不匹配问题,我们引入了块验证感知损失(BV损失),这是一种旨在最大化草稿序列期望接受长度的训练目标。BV损失直接由块验证接受规则推导而来,在草稿模型训练目标与序列级推理时验证机制之间建立了原则性的联系。在数学、代码和聊天基准测试中,对于使用Qwen3-4B和Qwen3-8B的DFlash和DSpark模型,BV损失在块验证下将每次验证调用平均接受的标记数比交叉熵损失训练提高了13.0%至21.0%,且不改变推理过程。BV损失也优于诸如TV损失和LK损失等逐标记接受目标,其增益扩展到标记验证和贪婪解码。这些结果证明了使用与序列级验证对齐的目标训练块扩散草稿模型(而非独立优化每个标记)的益处。
英文摘要
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0--21.0\% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.