发表机构
Department of Computer Science and Information Engineering, National Taiwan University(国立台湾大学计算机科学与资讯工程学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出DominoTree,一种无需训练的最佳优先草稿树,利用Domino的条件非因式分解校正评分。在Qwen3-4B等基准测试中,相比自回归解码加速显著,接受长度高,用GPU原生构建器保持低成本,在吞吐量上优于多种方法。
AI 中文摘要
推测性解码通过并行生成几个令牌并进行验证来加速大语言模型推理。像DFlash这样的块扩散草稿生成器一次生成一个草稿块,但仅对每个位置的边际概率进行建模;像DDTree这样的最佳优先树方法从这些边际概率扩展候选树。已发布的Domino草稿生成器添加了基于门控循环单元的因果校正,使每个草稿令牌的分布依赖于路径,这是DDTree的因式分解公式无法表示的结构。我们引入了DominoTree,一种无需训练的最佳优先草稿树,通过Domino的条件非因式分解校正沿着每条从根到节点的路径进行评分,通过将每个节点的校正限制为候选前M个来实现实用化。在八个基准测试中的Qwen3-4B上,DominoTree比自回归解码速度提高了6.6倍,并且在我们测试的每个温度下,达到了所有评估方法中最高的平均接受长度,每轮高达10.7个令牌。DominoTree使用GPU原生的CUDA图构建器构建其树,该构建器与参考Python实现位相同,因此接受度不变,同时保持每轮树构建成本低廉。以这个构建器为默认设置,DominoTree在每个温度下都比已发布的Domino解码器赢得了吞吐量,在Qwen3-4B上总体提高了9-10%,在Alpaca上高达22%,并且在我们测试的每个温度下都超过了DDTree/CaDDTree。在Qwen3-8B上,DominoTree在每个温度下都保持了最高的接受长度,并在T=0时决定性地赢得了吞吐量,比DDTree高24%;在较高温度下,它相对于DDTree/CaDDTree的优势缩小到平局和小损失,而其总体聚合优势超过DFlash和Domino仍然存在。
英文摘要
Speculative decoding accelerates LLM inference by drafting tokens and verifying them in parallel. Block-diffusion drafters such as DFlash model only per-position marginals, and tree methods such as DDTree expand candidate trees from those marginals. The released Domino drafter adds a GRU-based causal correction making each draft token's distribution path-dependent, a structure DDTree's factorized formulation cannot represent. We introduce DominoTree, a training-free best-first draft tree scored by Domino's conditional (non-factorized) correction along each root-to-node path, made practical by restricting the per-node correction to a candidate top-M. We evaluate it on eight benchmarks in a single-stream harness, and in SGLang, where it runs as an out-of-tree plugin against AR, DFlash, EAGLE-3 and Domino under identical flags. DominoTree attains the highest mean accepted length in every serving cell - two model sizes, single-request and concurrent load, context to 32K - and the highest Overall accepted length at every temperature in the research harness (21 of 24 per-dataset cells). A three-arm decomposition holding drafter, budget and verifier fixed separates the gain from applying the correction at all (+10.1% accepted length) from that of recomputing it along each candidate's realized path (+4.7% more), the part this paper adds. Where the round is verify-dominated, throughput follows: up to 7.3x over AR on Qwen3-8B, beating the released Domino decoder at its CUDA-graph best at every temperature, and inside SGLang winning single-request throughput by +12% over Domino on Qwen3-8B. On HELMET long context it beats Domino by +29-36% accepted length and +10-34% throughput at every length and both model sizes. Past a memory-constrained card's admission cap the chain wins goodput, and at our longest context, where prefill dominates, our lead over EAGLE-3 narrows to a tie.
Comments32 pages, 2 figures, 16 tables. Code: https://github.com/slin-zhq/Domino-Tree