arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24698cs.CLcs.AI

将树结构投机解码适配到DeepSeek-V4以实现高效推理

Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference

Changxu Liu, Zhaogeng Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对DeepSeek-V4的压缩注意力破坏跨分支状态一致性的问题,提出分支感知验证与状态隔离的树结构投机解码,在相同预算下提升接受长度和吞吐量,最高吞吐量提升约18.5%。

中文摘要 AI 辅助

在自回归解码过程中重复执行目标模型是导致大语言模型推理延迟的主要来源。与遵循单一候选链的线性投机不同,树结构投机从共享前缀保留多个分支;在相同预算下,这种更广泛的覆盖可以提高接受率和效率。将其适配到DeepSeek-V4并非易事:其CSA/HCA在线压缩注意力将困难集中在目标验证侧,其中从共享前缀分叉的分支压缩成不同状态,破坏了跨分支的状态一致性。我们通过分支感知的因果验证、临时状态隔离和接受路径状态刷新,将树结构投机解码集成到DeepSeek-V4-Flash流水线中,保持验证和压缩状态更新在各分支间的一致性。在预算D=5到D=8、批大小1到64以及三个数据集(GSM8K、MBPP、ShareGPT)上,树投机在所有设置中均实现了比匹配的线性配置更高的接受长度(例如,在D=8时约为2.83--3.41,而线性配置为2.39--2.84),并且在几乎所有配置中提高了吞吐量——仅在最小预算下提升微弱——最高提升约18.5%。更重要的是,这些收益遵循稳定且可迁移的规律:相对收益随预算增加而增长,并且在中小批大小下对可预测性较低的工作负载最为显著,而超过一定预算后吞吐量趋于平稳,并与仍在上升的接受长度脱钩。这些结果表明,在相同预算下保留多个候选路径可以有效提高DeepSeek-V4的解码效率,并为将投机解码适配到未来具有压缩、稀疏或结构化上下文表示的模型提供了经验。

英文摘要

Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which follows a single candidate chain, tree-structured speculation retains multiple branches from shared prefixes; under the same budget, this broader coverage can improve acceptance and efficiency. Adapting it to DeepSeek-V4 is nontrivial: its CSA/HCA online compressed attention concentrates the difficulty on the target-verify side, where branches diverging from a shared prefix compress into different states, breaking cross-branch state consistency. We integrate tree-structured speculative decoding into the DeepSeek-V4-Flash pipeline via branch-aware causal verification, temporary state isolation, and accepted-path state refresh, keeping verification and compressed-state updates consistent across branches. Across budgets D=5 to D=8, batch sizes 1 to 64, and three datasets (GSM8K, MBPP, ShareGPT), tree speculation achieves a higher accepted length than the matched linear configurations in all settings (e.g., at D=8 about 2.83--3.41 versus 2.39--2.84) and improves throughput in nearly all configurations---marginal only at the smallest budget---by up to about 18.5%. More importantly, the gains follow stable, transferable regularities: the relative gain grows with the budget and is most pronounced for less predictable workloads at small-to-medium batch sizes, while beyond a certain budget throughput plateaus and decouples from the still-rising accepted length. These results show that retaining multiple candidate paths under the same budget can effectively improve DeepSeek-V4 decoding efficiency, and offer experience for adapting speculative decoding to future models with compressed, sparse, or structured context representations.

发表机构

  • Baige AI Team, Baidu Inc.(百度百鸽AI团队)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑