arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过渐进式树状草稿的推测性解码解锁自回归语言模型中的并行性

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

Zipeng Gao, Zhi Zheng, Qingrong Xia, Junda Lin, Ziwei Zhao, Tong Xu, Zhefeng Wang, Enhong Chen

arXiv 2607.10661首次发表:更新:

发表机构

University of Science and Technology of China; Huawei Technologies Co., Ltd.(中国科学技术大学; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出渐进式树状草稿(PTD)方法,通过结构化、引导式并行草稿策略利用模型并行潜力,结合渐进树结构与逐步修剪机制,在单次前向传播中引导LLM探索多条语义路径,实现高达2倍解码加速且无需训练、与模型无关。

AI 中文摘要

推测性解码通过缓解内存受限瓶颈,显著加速了大语言模型(LLM)的推理。然而,传统的推测性解码通常依赖于辅助草稿模块,会产生大量的训练和通信开销。尽管最近的方法试图在目标模型本身内部生成草稿,但由于缺乏结构协调,它们往往无法充分利用其潜在的并行能力。在本文中,我们提出了渐进式树状草稿(PTD),它采用结构化、引导式并行草稿策略来利用模型的并行潜力。通过将渐进树结构与逐步修剪机制相结合,PTD在单次前向传播中积极引导LLM探索多条语义路径,确保草稿的多样性和连贯性。实验表明,PTD在各种基准测试中实现了高达2倍的解码加速,同时无需训练且与模型无关。我们的代码可在以下网址获取:此https网址。

英文摘要

Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks. However, traditional speculative decoding typically relies on auxiliary draft modules, incurring significant training and communication overhead. Although recent methods attempt to generate drafts within the target model itself, they often fail to fully exploit its latent parallel capacity due to a lack of structural coordination. In this paper, we propose \textbf{Progressive Tree Drafting (PTD)}, which employs a structured, guided parallel drafting strategy to harness the model's parallel potential. By coupling a progressive tree structure with a stepwise pruning mechanism, PTD actively guides the LLM to explore multiple semantic paths in a single forward pass, ensuring both draft diversity and coherence. Experiments demonstrate that PTD achieves up to $2\times$ decoding speedup across various benchmarks while remaining training-free and model-agnostic. Our code is available at: https://github.com/MINE-USTC/PTD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑