发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出期望解码轮数(EDR)目标,通过马尔可夫奖励过程直接优化推测解码的全局效率,无需辅助超参数,并在九个基准上提升草稿器性能。
AI 中文摘要
推测解码通过使用低成本的草稿模型提出令牌,由全尺寸目标模型并行验证,从而加速大型语言模型推理。并行和半自回归(semi-AR)草稿器通过在单次前向传播中提出整个块来提高起草效率,但训练它们带来了新的困难:给定位置的草稿分布取决于解码轮从哪里开始,而轮从哪里开始又取决于前几轮接受了多少令牌。现有的训练目标通常依赖于忽略这种跨轮耦合的块局部替代目标,因此不能直接优化全局解码效率。在这项工作中,我们通过将推测解码表示为马尔可夫奖励过程,开发了一个用于训练和评估这些草稿器的理论框架。该公式产生了期望解码轮数(EDR)目标,它通过状态占用率对局部拒绝成本进行加权,并精确等于期望解码轮数。与先前的替代目标不同,EDR不引入任何辅助超参数。然后,我们推导出一个精确的时间差分梯度,支持从目标模型 rollout 中进行无偏随机优化。同一框架还产生了一个用于轮数的精确离线评估器,使得在共享目标 rollout 上进行成对草稿器比较而无需运行推测解码成为可能。使用EDR对两个最先进的草稿器DSpark和DFly进行微调,一致地提高了平均接受长度,并在涵盖数学推理、代码生成和聊天的九个基准上优于现有训练目标。
英文摘要
Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process. This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds. Unlike prior surrogate objectives, EDR introduces no auxiliary hyperparameters. We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts. The same framework also yields an exact offline evaluator for round counts, enabling paired drafter comparisons on shared target rollouts without running speculative decoding. Finetuning two state-of-the- art drafters, DSpark and DFly, with EDR consistently improves mean accepted length and outperforms existing training objectives across nine benchmarks spanning math reasoning, code generation, and chat.