arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08721cs.CLcs.AI

LibraSpec:基于动态扩散的推测解码,通过边际增益驱动优化

LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization

  • Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州高等研究院)

机构由 AI 辅助整理,请以论文原文为准。

Zexun Lin, Yuan Feng, Junlin Lv, Kevin S. Zhou, Xike Xie

AI总结:

LibraSpec是一种无需训练的即插即用推测解码算法,通过边际增益驱动优化动态确定推测长度,可提升大语言模型推理速度,在多基准上较基线加速0.5~1.5倍、较自回归解码最高达8.49倍。

AI中文摘要:

推测解码通过生成多个token进行并行验证来加速大语言模型推理,其效率关键取决于每轮解码选择的推测长度。现有动态推测方法通过估计将被接受的token数量来选择推测长度,这对于按顺序生成token的自回归 draft 模型是合理的。然而,近期基于扩散的 draft 模型以显著更低的 draft 成本并行生成候选块,使关键问题从生成多少token转变为多少生成的token值得验证。因此,我们将动态推测长度选择重新表述为预期加速优化,并推导了一个边际准则:仅当推测序列的接受增益超过额外验证成本时才扩展该序列。基于此准则,我们开发了LibraSpec,这是一种无需训练、即插即用的算法,利用 draft 模型的置信度分数迭代确定推测长度。理论上,我们证明LibraSpec单调收敛到最优推测长度。在六个目标模型、三种基于扩散的推测解码方法以及数学、编码和聊天基准上的实验表明,在贪心和采样设置下均实现了一致的改进,较基线进一步实现0.5~1.5倍的加速,较自回归解码最高实现8.49倍的加速。

英文摘要:

Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing dynamic speculation methods select the speculation length by estimating how many tokens will be accepted, which is reasonable for autoregressive drafters that generates tokens sequentially. The recent wave of diffusion-based drafters, however, generates candidate blocks in parallel at substantially lower drafting cost, shifting the key question from how many tokens to generate to how many generated tokens are worth verifying. We therefore reformulate dynamic speculative-length selection as expected-speedup optimization and derive a marginal criterion that extends the speculative sequence only when its acceptance gain outweighs the additional verification cost. Building on this criterion, we develop \textit{LibraSpec}, a training-free and plug-and-play algorithm that iteratively determines the speculative length using drafter confidence scores. Theoretically, we prove that LibraSpec monotonically converges toward the optimal speculative length. Experiments across six target models, three diffusion-based speculative decoding methods, and math, coding, and chat benchmarks show consistent improvements under both greedy and sampling settings, achieving a further $0.5\sim1.5\times$ improvement over baselines and up to $8.49\times$ speedup over autoregressive decoding.

↑