arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaFlash:通过策略蒸馏扩散草稿器实现自适应推测解码

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

Yu-Yang Qian, Hao-Cong Wu, Chen Chen, Jiacheng Sun, Zhenhua Dong, Peng Zhao, Zhi-Hua Zhou

arXiv 2607.19223首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University; Huawei Foundation Model Dept(南京大学计算机软件新技术国家重点实验室; 南京大学人工智能学院; 华为基础模型部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散草稿器在推测解码中存在的问题,提出AdaFlash框架,包含基于策略蒸馏算法和自适应长度头,有效降低方差,实验证明该框架能提高推理速度,尤其在高并发场景中吞吐量显著提升。

AI 中文摘要

推测解码是一种加速大语言模型推理的流行范式,其中轻量级草稿模型先生成草稿序列,再由目标模型并行验证。DFlash等工作利用扩散草稿器进一步提高了草稿生成效率。但本文发现扩散草稿器存在双向注意力这一双刃剑问题,它虽赋予模型并行生成和全局上下文建模能力,但也在域级和令牌级引入高方差。为此提出AdaFlash框架,包括为扩散草稿器定制的基于策略蒸馏算法和自适应长度头,实验表明该框架在部署时持续提高加速率,在高并发场景中吞吐量比之前方法高出约66%。

英文摘要

Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass. In this work, we uncover a central pitfall of diffusion drafters: the global dependency, arising from bidirectional attention and KV injection in diffusion drafters, is a double-edged sword. On one hand, it endows the model with parallel generation and global contextual modeling capabilities; on the other hand, this global dependency introduces high variance at both the domain-level and the token-level: acceptance rates fluctuate substantially across different domains, and the acceptance probability varies across draft positions. To tackle this issue, we propose the AdaFlash framework, comprising two components: (i) an on-policy distillation (OPD) algorithm with reverse-KL divergence tailored for diffusion drafters, which delivers stable convergence and continuously adapts the drafter to the deployment distribution; and (ii) an adaptive length head that dynamically adjusts the candidate length on the fly, substantially lowering the verification cost of the target model. Experiments demonstrate that AdaFlash consistently improves the speedup during deployment, with especially significant gains under high-concurrency, achieving up to 66% higher average throughput than previous state-of-the-art methods. Our code is available at https://github.com/ZinYY/AdaFlash.

CommentsCOLM'26 Workshop (Spotlight)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑