发表机构
Independent Researcher; Baidu Inc.; Zhejiang University(独立研究者; 百度公司; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分析有损推测解码的机制与失效模式,将其分为两类并构建评估框架,发现基于截断的有损方法会因分布失真致性能下降,协作类需控制概率超调以保障生成质量。
AI 中文摘要
推测解码(Speculative Decoding, SD)通过轻量草稿模型提出token、再由更大的目标模型并行验证,加速大语言模型推理。近期方法引入有损验证方案,通过放宽严格分布匹配进一步提升效率,但这种放宽会悄然改写解码分布,带来的加速可能以生成质量不稳定、甚至严重下降为代价。本研究对有损验证方法诱导的分布进行了原则性分析,发现诸多看似不同的方法仅表面有别,可分为基于截断的验证与协作验证两类;在精心构建的基准上搭建诊断评估框架后,确定基于截断的方法存在根本缺陷:因分布失真,性能较真实截断采样基线显著下降;针对协作验证,揭示了关键原则:控制草稿概率相对目标概率的超调是防止低质量输出的核心。本研究代码可在指定网址获取。
英文摘要
Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verified in parallel by a larger target model. Recent approaches introduce lossy verification schemes to further improve efficiency by relaxing strict distributional matching. Yet such relaxation silently rewrites the decoding distribution, and the resulting acceleration can come at the cost of unstable, sometimes severely degraded generation quality. In this work, we present a principled analysis of the distributions induced by lossy verification methods. We show that many seemingly distinct approaches differ only superficially and can be unified into two categories: truncation-based verification and collaborative verification. We further construct a diagnostic evaluation framework across curated benchmarks. For truncation-based methods, we identify a fundamental pitfall-performance can degrade significantly compared to the true truncation sampling baseline due to distributional distortion. For collaborative verification, we reveal that well-designed relaxation principles, namely overshoot suppression and supervision quality, matter far more than the linear interpolation between draft and target. Our code is available at https://github.com/ZhouYuxuanYX/Fast-HSD.