发表机构
Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对高并发下推测解码中长草稿浪费计算资源问题,提出D-Cut自适应剪枝方法,跨批次联合选草稿令牌,依接受长度差异和运行时成本模型自适应调整,实验表明该方法能显著提升加速比。
AI 中文摘要
推测解码可在不影响输出质量的情况下加速大语言模型推理。近期的并行起草方法通过将草稿长度与起草延迟解耦进一步提升单请求性能。然而,在高请求并发下,长草稿会在被拒绝的令牌上浪费大量计算,增加验证成本。我们提出D-Cut,一种自适应剪枝方法,它跨批次联合选择草稿令牌,并将验证预算集中在最有可能被接受的令牌上。D-Cut基于两个观察结果:并发请求的接受长度差异很大,因此它进行跨请求剪枝;验证成本强烈依赖于部署环境,所以它纳入运行时成本模型以使其剪枝深度适应目标环境。在密集模型和专家混合模型上的实验表明,在高并发下,D-Cut将平均加速比从1.26倍提高到1.65倍,在长草稿基线比自回归解码慢的密集模型配置中恢复加速,并在专家混合模型上比自回归解码实现高达3.0倍的加速。
英文摘要
Speculative decoding accelerates large language model (LLM) inference without compromising output quality. Recent parallel drafting methods further improve single-request performance by decoupling draft length from drafting latency, enabling longer drafts and higher mean accepted tokens (MAT). However, under high request concurrency, long drafts waste substantial computation on rejected tokens, increasing verification cost and potentially making speculative decoding slower than autoregressive decoding. We present D-Cut, an adaptive pruning method that selects draft tokens jointly across the batch and concentrates the verification budget on tokens most likely to be accepted. D-Cut is motivated by two observations. First, acceptance lengths vary considerably across concurrent requests; D-Cut therefore performs cross-request pruning, allocating the verification budget adaptively according to draft confidence. Second, verification cost depends strongly on the deployment environment, including GPU architecture and parallelism strategy; D-Cut incorporates a runtime cost model to adapt its pruning depth to the target environment. Experiments on dense and mixture-of-experts (MoE) models show that, under high concurrency, D-Cut improves the average speedup from \(1.26\times\) to \(1.65\times\), restores acceleration in dense-model configurations where long-draft baselines are slower than autoregressive decoding, and achieves up to \(3.0\times\) speedup over autoregressive decoding on MoE models.