arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并行起草,深度条件化:用于投机解码的相邻因果注入

Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

Haohui Zhang, Keyu Chen, Haocheng Sun, Weibo Gu, Ruizhi Qiao, Xing Sun, Bo Jiang

arXiv 2609.36173首次发表:更新:

发表机构

Shanghai Jiao Tong University; Tencent YouTu Lab; Xiamen University(上海交通大学; 腾讯优图实验室; 厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DSpine,通过主干网络逐层门控相邻因果注入实现并行起草与深度条件化,在多个基准上超越DFlash,显著提升投机解码的接受长度和吞吐量。

AI 中文摘要

并行投机起草在一次主干网络前向传播中生成多个候选,但独立的词元选择可能产生不一致的延续,从而缩短被接受的先前片段。现有方法大多将条件解码留给主干网络之后的轻量级模块,这限制了前驱信息向后继的流动。我们对DFlash的分析表明,早期位置在浅层已经形成可恢复的预测,且准确的相邻前驱越早进入越能帮助后继。因此,我们提出DSpine,一种在主干网络全程进行因果条件注入的起草器:在每一层,门控相邻注入将每个前驱的预测特征写入其后继,使得因果条件链在网络深度上展开,同时所有位置并行更新。基于目标模型输出嵌入构建的统一转移空间将逐层注入与前驱条件化解码统一起来,且逐层输出嵌入监督促进浅层预测特征的形成。融合内核和转移缓存使得两者在SGLang内高效并行执行。在七个数学、代码和聊天基准上,DSpine在Qwen3-4B和Qwen3-8B上均取得最长接受长度。在Qwen3-8B零温度下,它将七个基准的平均值从DFlash的3.77提升至4.82(+27.8%);在SGLang服务测试中,它平均比DFlash提供23.3%更高的吞吐量。

英文摘要

Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor's predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model's output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B. At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑