arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边际而非窗口:无训练的逐步有损推测解码

Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

Oszkár Urbán, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo

arXiv 2609.02897首次发表:更新:

发表机构

University of Cambridge; Samsung AI Center-Cambridge, UK(剑桥大学; 英国三星剑桥人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无训练的逐步推测解码方法AdaptiveSpec,通过调整token匹配规则与draft树形状,在SGLang引擎上实现比EAGLE-3最高56%的吞吐量提升,同时保持93%至完全无损的任务准确率。

AI 中文摘要

推测解码通过生成候选token并并行验证来加速大语言模型(LLM)推理,EAGLE-3等树注意力 draft 生成器被广泛采用,但通常固定两个决策:(1)严格的token匹配验证规则,(2)静态的 draft 树形状。现有工作在有限假设下分别放松这两点:无训练有损验证需长 draft 链,固定token预算下采用自适应树形状。本文提出AdaptiveSpec,一种无训练的逐步推测解码方法,可从解码过程中已产生的内部信号调整上述两个决策。逐步边际规则:当目标token的概率与 draft 提出token的概率之比超过阈值时,允许不匹配的 draft 提出token,无需依赖 draft 长度或底层 draft 生成器架构;逐步树策略:从 draft top-1置信度与捕捉近期 draft-目标一致性的滚动接受历史的融合信号,直接调整 draft 树的深度、宽度和节点数,允许总 draft 数量变化而非仅重新分配。两种调整在正交轴上运行,效果叠加。在生产级服务引擎SGLang上实现后,AdaptiveSpec比当前最先进的自回归推测解码方法EAGLE-3的吞吐量提升最高达56%,在三个目标模型(DeepSeek-R1-Distill-Llama-8B、Llama-3.1-8B-Instruct、Qwen3-8B)的GSM8K、MATH-500和HumanEval任务上恢复了93%至完全无损的任务准确率。

英文摘要

Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict token-match verification rule and (2) a static draft-tree shape. Prior work relaxes each in isolation under limiting assumptions: long draft chains for training-free lossy verification, and adaptive tree shaping under a fixed token budget. We introduce AdaptiveSpec, a training-free per-step speculative decoding method that adapts both decisions from internal signals already produced during decoding. A per-step margin rule promotes a mismatched draft-proposed token when the ratio of the target's probability on the drafted token to its top-1 probability exceeds a threshold with no dependence on draft length or underlying drafter architecture. A per-step tree policy adjusts the draft tree's depth, width, and node count directly from a fused signal of draft top-1 confidence and a rolling acceptance history capturing recent draft-target agreement, allowing the total draft count to vary rather than only be redistributed. The two adaptations operate on orthogonal axes and compound in effect. Implemented on the SGLang production-grade serving engine, AdaptiveSpec improves throughput over the state-of-the-art autoregressive speculative decoding method EAGLE-3 by up to 56%, recovering 93% to fully lossless task accuracy across GSM8K, MATH-500, and HumanEval on three target models (DeepSeek-R1-Distill-Llama-8B, Llama-3.1-8B-Instruct, Qwen3-8B).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑