arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24411cs.AI

ResiSpec:通过残差分布塑形增强多候选推测采样

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye

首次发表
浏览论文内容

中文总结 AI 辅助

ResiSpec 框架针对多候选推测采样的残差漂移问题,通过重塑验证阶段的提议分布,实现最高1.92倍的加速,提升了LLM服务效率。

中文摘要 AI 辅助

大型语言模型(LLM)服务的效率本质上受限于自回归解码的顺序性。推测解码(SD)通过使用轻量草稿模型推测未来 token,再由 LLM 在单次并行前向传播中验证这些 token,以此缓解该问题。为进一步提升效率,多候选方案提出多样化候选集以提高 token 被接受的概率。然而,本文表明这些方案受残差漂移瓶颈制约:初始候选被拒绝会导致残差目标分布偏离草稿模型的预测,该偏移使后续候选失效,迫使系统进行代价高昂的重采样。为解决此问题,本文提出 ResiSpec 框架,其在验证过程中策略性地重塑提议分布,将残差目标质量锚定在草稿模型的高置信度区域内。ResiSpec 通过数学上重新对齐验证过程且不损害输出准确性,防止候选失效,实现了比最先进多候选方法高至 1.92 倍的加速。代码可在该 https URL 获取。

英文摘要

The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of autoregressive decoding. Speculative Decoding (SD) mitigates this by using a lightweight draft model to speculate future tokens, which are then validated by the LLM in a single parallel forward pass. To further boost efficiency, multi-candidate schemes propose diverse candidate sets to increase the likelihood of token acceptance. However, we show that these schemes are bottlenecked by Residual Drift: a phenomenon where the rejection of initial candidates causes the residual target distribution to diverge from the draft model's predictions. This shift renders subsequent candidates ineffective and forces the system into expensive resampling. To resolve this, we propose ResiSpec, a framework that strategically reforms the proposal distribution during verification to anchor the residual target mass within the draft model's high-confidence regions. By mathematically re-aligning the verification process without compromising output exactness, ResiSpec prevents candidate obsolescence and achieves up to 1.92$\times$ speedup over state-of-the-art multi-candidate methods. Code is available at https://github.com/Czzzk/Resispec.

发表机构

  • School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
  • National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑