无损但不免费:消费级硬件上推测性解码的实证剖析
Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
浏览论文内容
中文总结 AI 辅助
研究大语言模型单流自回归解码受内存带宽限制问题,核心方法是推测性解码,通过小草稿模型与目标模型配合及拒绝采样规则重构计算。在苹果硅笔记本上实证研究五种配置,验证分布等效性,揭示加速与减速原因,表明该方法在特定条件下才有效。
中文摘要 AI 辅助
大语言模型的单流自回归解码受内存带宽限制,每次生成令牌都需完整前向传递目标模型,且连续传递无法并行化。推测性解码重构了这种计算,小草稿模型自回归提出K个令牌,目标模型批量评分,拒绝采样规则可保持目标模型输出分布。本文给出从零开始、设备无关的实现,并在苹果硅笔记本上对五种草稿/目标后端配置进行实证研究。在三个层面验证了分布等效性,最佳配置在K = 6时实现了1.61倍的壁钟加速,同时也揭示了部分配置减速的原因。结果表明,推测性解码只有在验证真正批并行且草稿/目标延迟差距真实时才会有成效。
英文摘要
Single-stream autoregressive decoding of large language models is bound by memory bandwidth: each generated token requires one full forward pass through the target model, and successive passes cannot be parallelized. Speculative decoding restructures this computation: a small draft model proposes $K$ tokens autoregressively, the target model scores all of them in one batched pass, and a rejection-sampling rule provably preserves the target model's output distribution. We present a from-scratch, device-agnostic (CUDA/MPS/CPU) implementation and an empirical study across five draft/target backend configurations on a consumer Apple-silicon laptop. Distribution equivalence is verified at three levels, culminating in a two-sample test over roughly 9,200 real-model tokens per method ($χ^2 = 162.5$, dof $= 200$, $p = 0.976$) and exact greedy-sequence agreement. The best configuration reaches a measured $1.61\times$ wall-clock speedup at $K=6$, on an acceptance profile declining from 69.7% at $K=1$ to 37.8% at the optimum, while three of five configurations decelerate, either because the draft fails to out-speed a small target or because the quantized Metal backend executes "parallel" verification serially, an effect we isolate and quantify. The failures are as instructive as the successes: speculative decoding pays off only when verification is genuinely batch-parallel and the draft/target latency gap is real.
发表机构
- University of California, San Diego(加利福尼亚大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。