arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23900cs.LG

GDN Tree-Scan:递归混合语言模型的已服务树验证

LumoTree: Path-Parallel Speculative Verification for Hybrid Language Models

Zhiyuan Ma

AI总结:

提出GDN Tree-Scan,一种集成于vLLM的Gated-DeltaNet混合语言模型树验证器,通过分支局部状态扫描等技术,在Qwen3.6-27B上实现27.0%的令牌加权解码吞吐量提升。

AI中文摘要:

树形投机解码在一次目标前向传播中验证多个候选续写。对于仅注意力的Transformer,验证器主要需要一个祖先掩码。递归混合语言模型打破了这一假设:候选行还必须携带原生顺序解码沿其根到节点路径本应产生的递归状态。否则,验证器可以使用正确的注意力掩码,但仍基于不可能的递归历史进行条件化。我们提出了GDN Tree-Scan,一个集成到vLLM中的Gated-DeltaNet混合语言模型的已服务验证器。该系统结合了FlashAttention-2树偏置注意力、分支局部GDN扫描/重放、设备端多草稿提交以及仅接受链状态发布。在公开的Qwen3.6-27B-FP8检查点上,在温度0.6的干净批大小一(B=1)SWE/Codex解码门控中,一个六节点根分支树在接近原生的验证前向时间内将提交的令牌/事件提高了17.2%,并达到23.88令牌加权解码令牌/秒,而原生五步MTP(E5)为18.80,即27.0%的令牌加权解码吞吐量增益。每请求等延迟视角为+4.0%,端到端任务墙钟时间仍以前缀填充为主。经验等价性证据仅限于在观察到的原生翻转下限内的递归预言机概率重评分(p-rescore)闭包,而非完整的分布距离证明。

英文摘要:

Tree speculative decoding for recurrent-hybrid language models requires each accepted path to maintain consistent recurrent state, convolution history, and attention caches. We present LumoTree, a GPU serving design that coordinates verification and commitment through a shared tree descriptor. The verifier processes independent paths in parallel and reuses recurrent state tiles across local updates. Boundary states connect dependent paths. Accepted-path replay publishes the selected continuation into native running state, alongside coordinated convolution gathering and attention-cache remapping. Fused selection, device-resident acceptance, graph replay, and tree-aware attention integrate this organization into the serving cycle. The design separates temporary branch computation from persistent request state while retaining native prefix-cache interfaces. We evaluate numerical behavior, serving performance, and coding-agent outcomes on NVIDIA DGX Spark.

补充信息

↑