发表机构
Bahçeşehir University(巴赫切希尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于PEFT - BD的推测性解码方法,虽有避免分词器不匹配等优点,但在Qwen3 - 0.6B实验中未实现实际加速,因草稿生成器计算不高效。结果表明成功推测性解码需草稿生成器执行成本远低于验证器。
AI 中文摘要
推测性解码通过使用廉价的草稿生成器提出多个未来令牌并由目标模型进行验证来加速自回归语言模型推理。一个常见的设计目标是在减少辅助参数和系统开销的同时提高草稿质量。我们通过PEFT - BD研究了这个方向的负面结果,它是一种同骨干推测性解码方法,其中类似LoRA的适配器充当自回归验证器的块扩散草稿生成器。尽管它有避免分词器不匹配、避免加载单独草稿模型、仅添加少量可训练参数等优点,但在Qwen3 - 0.6B实验中未实现实际加速。虽然该方法获得了可观的接受前缀,但分析表明每个推测步骤都需要启用适配器的全骨干草稿传递,然后是禁用适配器的全骨干验证传递。因此,草稿生成器在参数上高效但在计算上不高效。我们的结果确定了成功推测性解码的一个简单但重要的条件:草稿生成器的执行成本必须远低于验证器。当草稿计算仍为验证器规模时,仅更长的接受前缀无法弥补。
英文摘要
Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a target model to verify them. A common design goal is therefore to improve draft quality while reducing auxiliary parameters and systems overhead. We study a negative result for this direction through PEFT-BD, a same-backbone speculative decoding method in which a LoRA-like adapter acts as a block-diffusion drafter for an autoregressive verifier. PEFT-BD is motivated by several attractive properties: it avoids tokenizer mismatch, avoids loading a separate draft model, adds only a small number of trainable parameters, and uses a BD3LM-style denoising objective to propose a block of tokens in parallel. Despite these advantages, PEFT-BD does not yield a practical speedup in our Qwen3-0.6B experiments. Although the method obtains nontrivial accepted prefixes, profiling shows that each speculative step requires an adapter-enabled full-backbone draft pass followed by an adapter-disabled full-backbone verification pass. Thus, the drafter is parameter-efficient but not compute-efficient. Our results isolate a simple but important condition for successful speculative decoding: the drafter must be substantially cheaper to execute than the verifier. Longer accepted prefixes alone cannot compensate when draft computation remains verifier-scale.