AI 中文总结
研究自回归语言模型中推测性解码的局限性,提出用Weaver适配器从因式草稿生成器边际构建提议树,恢复条件依赖,还推导算法并实现内核,结合模型与系统贡献,相比自回归解码加速4.37倍,性能超DFlash基线24.7%。
AI 中文摘要
推测性解码通过在单次前向传递中用计算换取额外生成的令牌,极大地提高了自回归语言模型的交互性。因式草稿模型特别高效,因为它们并行预测未来令牌的边际,但随着推测预算的增加,其独立性假设会导致接受率急剧下降。我们分析了这一限制,并引入了Weaver,这是一种轻量级自回归适配器,它从因式草稿生成器的前K个边际构建提议树。Weaver恢复了提议令牌之间的条件依赖关系,同时避免了全词汇投影。为了支持对具有门控Delta Net层的模型进行快速验证,我们推导了一种无回滚树验证算法,并在SGLang中实现了优化的CUDA内核。通过结合这些模型和系统贡献,我们比自回归解码实现了4.37倍的加速,并且比高度优化的DFlash基线性能高出24.7%。
英文摘要
Speculative decoding greatly increases the interactivity of autoregressive language models by trading off computation for extra tokens generated in a single forward pass. Factorized draft models are especially efficient because they predict future-token marginals in parallel, but their independence assumption causes acceptance rates to degrade sharply as the speculative budget grows. We analyze this limitation and introduce Weaver, a lightweight autoregressive adapter that constructs proposal trees from the top-K marginals of a factorized drafter. Weaver restores conditional dependencies between proposed tokens while avoiding a full-vocabulary projection. To support fast verification for models with Gated Delta Net layers, we derive a rollback-free tree-verification algorithm and implement optimized CUDA kernels in SGLang. By combining these model and systems contributions we achieve a 4.37-fold speedup over autoregressive decoding, and outperform a highly optimized DFlash baseline by 24.7%.