发表机构
The University of Tokyo; National Center for Mathematics and Interdisciplinary Sciences (NCMIS), AMSS, CAS(东京大学; 中国科学院数学与系统科学研究院数学与交叉科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Tsubame提出两遍树投机解码框架,通过先规划拓扑再重放采样,结合扩散起草器,提升解码接受长度与吞吐量。
AI 中文摘要
上下文感知的动态树根据草稿路径概率分配投机解码预算,使其深度和分支适应当前上下文。然而,在随机解码下,我们发现这种结构优势并不总能弥补随机采样与高级验证相结合的接受度提升,在某些设置下,动态树可能落后于采样链。这些树从候选本身生长其拓扑结构,因此提交验证的令牌通常是构建期间选择的确定性高得分令牌。这种耦合并非固有:一旦拓扑固定,其节点可以通过采样重新填充,使动态树在保留结构优势的同时,也能受益于随机采样和高级验证。基于扩散的起草器使这变得可行,因为它们的并行输出或轻量级条件修正允许在完整拓扑已知后廉价地重新生成候选。我们引入了Tsubame,一个用于基于扩散的起草器的两遍树投机解码框架。第一遍使用草稿路径分数规划并冻结上下文感知拓扑;第二遍重放固定拓扑,采样填充其节点的令牌以形成用于验证的候选树。我们证明Tsubame在兼容的采样和验证策略下是无损的。在三个基于扩散的起草器、六个数据集和多个候选预算上的实验表明,Tsubame在确定性树上提高了接受长度和吞吐量,包括逆转其对采样链劣势的设置。
英文摘要
Context-aware dynamic trees allocate the speculative decoding budget according to draft path probabilities, adapting their depth and branching to the current context. Under stochastic decoding, however, we find that this structural advantage does not always compensate for the acceptance gains of random sampling paired with advanced verification, and such dynamic trees can fall behind sampled chains in some settings. These trees grow their topology from the candidates themselves, so the tokens submitted for verification are typically the deterministic high-score tokens selected during construction. This coupling is not inherent: once the topology is fixed, its nodes can be repopulated by sampling, allowing dynamic trees to retain their structural advantage while also benefiting from random sampling and advanced verification. Diffusion-based drafters make this practical, as their parallel outputs or lightweight conditional corrections allow candidates to be regenerated cheaply after the complete topology is known. We introduce Tsubame, a two-pass tree speculative decoding framework for diffusion-based drafters. The first pass plans and freezes a context-aware topology using draft path scores; the second replays the fixed topology, sampling the tokens that populate its nodes to form the candidate tree for verification. We prove that Tsubame is lossless under compatible sampling and verification strategies. Experiments across three diffusion-based drafters, six datasets, and multiple candidate budgets show that Tsubame improves acceptance length and throughput over deterministic trees, including settings where it reverses their disadvantage against sampled chains.