发表机构
University of California, Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM推理延迟问题,提出无需训练的SSR自推测解码方法,结合思维链与后缀解码,在Qwen3.5、Gemma-4等模型上实现总生成延迟最高24.1%的相对提升。
AI 中文摘要
大型语言模型(LLMs)被部署用于越来越复杂的涉及规划和多步骤决策的任务,但这些任务的高质量性能通常需要生成长的推理链。这与对延迟敏感的交互式应用(如语音助手或编码智能体)不匹配,在这些应用中,生成长度延迟会严重影响用户体验。现有的加速方法通常专注于 token 级别的生成,而未利用推理工作流的结构。我们引入 SSR:用于推理模型的自推测技术(Self-Speculation for Reasoning Models),这是一种无需训练的自推测解码方法,利用思维链(Chain-of-Thought,CoT)作为推测源。SSR 将部分思维链(partial-CoT)的答案分布用作草稿器,将完整思维链(full-CoT)的分布用作验证器,两者均来自同一模型但使用不同的推理预算。这基于一个观察:后续的部分思维链响应通常与全预算响应具有更大的语义和词汇重叠。由于这种重叠,SSR 可以一次性接受长的草稿前缀,从而在结构化和长文本生成任务上实现大幅加速。为了进一步利用超出标准推测解码所接受的连续前缀的草稿-响应重叠,SSR 还结合了后缀解码,利用草稿生成后缀缓存并恢复超出接受前缀的有用片段,从而在草稿与最终响应具有高词汇重叠的任务上进一步降低延迟。我们在多个最适合 SSR 的结构化和长文本生成任务上对其进行评估,结果表明,对于 Qwen3.5 和 Gemma-4 等流行开源模型,SSR 在总生成延迟上实现了高达 24.1% 的相对提升。
英文摘要
Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor fit for latency-sensitive and interactive applications like voice assistants or coding agents, where generation latency can strongly affect user experience. Existing acceleration methods typically focus on token-level generation, without utilizing the structure of reasoning workflows. We introduce SSR: Self-Speculation for Reasoning Models, a training-free self-speculative decoding method that leverages the chain-of-thought (CoT) as a source of speculation. SSR uses the partial-CoT answer distribution as the drafter and the full-CoT distribution as the verifier, deriving both from the same model at different reasoning budgets. This builds on the observation that later partial-CoT responses often exhibit greater semantic and lexical overlap with the full-budget response. Due to this overlap, SSR can accept long draft prefixes at once, leading to large speedups on structured and long-form generation tasks. To further exploit draft-response overlap beyond the contiguous prefix accepted by standard speculative decoding, SSR also incorporates suffix decoding, using the draft to seed a suffix cache and recover useful spans beyond the accepted prefix, further reducing latency on tasks with high lexical overlap between the draft and the final response. We evaluate SSR on multiple structured and long-form generation tasks where it is most useful, and demonstrate a relative improvement of up to 24.1% on total generation latency for popular open-source models such as Qwen3.5 and Gemma-4.