发表机构
UNIST(蔚山科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SharpDraft通过基数感知查询缩放解决长上下文投机解码中的注意力质量稀释问题,无需训练即可在多个基准上实现2.59-3.19倍的端到端加速。
AI 中文摘要
长形式推理使得推理过程代价高昂,而投机解码通过并行验证多个草稿令牌来缓解这一成本。然而,随着上下文增长和草稿接受率下降,其加速效果可能会减弱。我们关注注意力质量稀释问题:当softmax对更多可见的键进行归一化时,集中在最高得分键上的质量可能会减少。我们提出了SharpDraft,一种无需训练的方法,通过基数感知的查询缩放来抵消这一效应,而无需在线自适应的计算开销。在明确假设下,我们推导出精确的top-$k$质量修正,并部署了闭式固定斜率近似。在AIME-26、GPQA-Diamond和LongGenBench Diary上,SharpDraft在应用于DFlash、PARD和EAGLE 3.1时,相较于仅目标的自回归解码,实现了$2.59$-$3.19\ imes$的几何平均端到端加速。与DFlash结合时,它提高了解码速度,并在端到端加速上优于全参数和基于LoRA的在线自适应,同时匹配未修改草稿器报告的峰值分配GPU内存。
英文摘要
Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentrated on the highest-scoring Keys can decrease. We introduce SharpDraft, a training-free method that counteracts this effect through cardinality-aware Query scaling, without the computational overhead of online adaptation. Under explicit assumptions, we derive an exact top-$k$ mass correction and deploy a closed-form fixed-slope approximation. Across AIME-26, GPQA-Diamond, and LongGenBench Diary, SharpDraft achieves $2.59$-$3.19\times$ geometric-mean end-to-end speedups over target-only autoregressive decoding when applied to DFlash, PARD, and EAGLE 3.1. With DFlash, it improves decoding speed and outperforms full-parameter and LoRA-based online adaptation in end-to-end speedup, while matching the unmodified drafter's reported peak allocated GPU memory.
Comments19 pages, 7 figures