LongSpark:具有固定成本并行起草器的高效投机解码
LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
浏览论文内容
中文总结 AI 辅助
LongSpark提出一种固定成本并行起草器,通过从目标验证中提取多尺度视图,实现解码成本与上下文长度无关,在长上下文任务中达到最优效率并大幅减少状态。
中文摘要 AI 辅助
投机解码通过在单次目标前向传播中验证多个草稿令牌来加速自回归推理。然而,随着上下文增长,现有的最先进起草器变得日益昂贵,削弱了它们本应提供的效率优势。我们认为这种扩展是不必要的。独立语言模型必须随其前缀增长,因为它对生成的每个令牌全权负责。相比之下,起草器仅提出候选;目标在提交任何令牌之前捕获并纠正每个错误。因此,起草器的解码成本可以完全独立于前缀长度。我们提出了LongSpark,一种块扩散起草器,通过从目标的验证过程中提取固定大小、多尺度的视图来实现这一目标,从而消除了对增长持久状态的需求。大量评估表明,LongSpark在多种模型规模和现实服务条件下实现了最先进的端到端效率。值得注意的是,它在长上下文任务中提供了最低的每输出令牌时间,同时将起草器的上下文状态减少了几个数量级。
英文摘要
Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting fixed-size, multiscale views from the target's verification pass, thereby eliminating the need for a growing persistent state. Extensive evaluations demonstrate that LongSpark achieves state-of-the-art end-to-end efficiency across multiple model scales and realistic serving conditions. Notably, it delivers the lowest time-per-output-token on long-context tasks while reducing the drafter's context state by several orders of magnitude.