核采样投机解码:超越精确分布的合理性感知验证
Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution
浏览论文内容
中文总结 AI 辅助
针对投机解码中标准接受规则过于保守的问题,提出核采样投机解码(NSD),通过将目标模型核纳入接受条件实现宽松验证,显著提升解码效率并保持任务性能。
中文摘要 AI 辅助
投机解码通过使用轻量级草稿模型提出多个令牌,并由目标模型并行验证,从而加速自回归生成。然而,标准接受规则侧重于精确分布校正,当草稿模型分配过多概率时,会拒绝在目标模型下仍然高度合理的令牌。这种保守的验证限制了每次验证前向传递后保留的草稿令牌数量。我们引入了核采样投机解码(NSD),一种将目标模型合理性纳入投机解码的宽松验证方法。NSD接受一个草稿令牌,如果它满足标准接受规则或属于目标模型的核(nucleus)。我们从理论上表征了该方法引入的分布偏差,并表明单步误差完全由草稿模型在目标核内的多余概率决定。我们进一步推导了序列级保真度界限,量化了局部偏差在自回归解码过程中如何累积。跨多个目标模型和提议机制的实验表明,NSD持续提高投机解码效率,同时保持具有竞争力的任务性能。我们的方法在自回归解码上实现了高达5.16倍的吞吐量加速,在标准投机解码上实现了高达3.15倍的加速。这些改进与更长的接受长度相吻合,使得更多输出令牌能够分摊每次目标验证前向传递的成本。分析表明,合理性感知验证为宽松验证和投机解码效率提供了一种有效方法。我们的代码可在https URL获取。
英文摘要
Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel. However, the standard acceptance rule focuses on exact distribution correction and rejects tokens that remain highly plausible under the target model when the draft model assigns excess probability. This conservative verification limits the number of draft tokens retained after each verification forward pass. We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding. NSD accepts a draft token if it satisfies the standard acceptance rule or belongs to the target model's nucleus. We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model's excess probability within the target nucleus. We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding. Experiments across multiple target models and proposal mechanisms demonstrate that NSD consistently improves speculative decoding efficiency while maintaining competitive task performance. Our method achieves throughput speedups of up to $5.16\times$ over autoregressive decoding and up to $3.15\times$ over standard speculative decoding. These improvements coincide with longer accepted lengths, allowing more output tokens to share the cost of each target verification pass. Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency. Our code is available at https://github.com/EIT-NLP/Nucleus-Speculative-Decoding.
发表机构
- Eastern Institute of Technology(东方理工大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。