TreeWY:门控DeltaNet混合模型的推测验证
TreeWY: Speculative Verification for Gated DeltaNet Hybrids
浏览论文内容
中文总结 AI 辅助
TreeWY通过树结构WY变换移除GDN层的循环状态快照,降低混合模型推测解码的内存压力,提升吞吐量与TTFT,还支持更宽的高接受度draft。
中文摘要 AI 辅助
现代开源模型多为混合架构:多数层是门控DeltaNet(GDN)层,携带固定大小的循环状态而非不断增长的键值(KV)缓存,这使普通解码内存效率高,但会损害推测解码。为验证一批 draft token 并回退被拒绝的 token,当前系统会在每个 draft 位置为 GDN 层快照完整循环状态,且这些快照无法在 draft 树的分支间共享,因此宽的高接受度树会因内存不足而不可行。我们移除了这些快照,通过门控 delta 规则的树结构 WY 变换,用单次三角求解计算每个 draft 节点的输出,并在提交时仅重建被接受的状态,存储小型伪值矩阵而非每个节点的状态;该推导仅依赖门控 delta 规则,不依赖其他架构细节。在两个规模的混合模型系列(Qwen3.5 35B 和 397B)的服务基准测试中,这在相同接受长度下降低了推测循环状态内存和 KV 缓存压力,将释放的高带宽内存(HBM)转化为更高吞吐量和低得多的首 token 时间(TTFT),在内存受限处效果显著,在非内存受限处仅增加少量开销。相同内存还提升了树宽:可实现更宽、高接受度的 draft,不过尚未带来吞吐量增益。
英文摘要
Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache. This makes ordinary decoding memory-efficient, but hurts speculative decoding. To verify a batch of draft tokens and then roll back the rejected ones, today's systems snapshot the full recurrent state at every draft position for GDN layers, and those snapshots cannot be shared across branches of a draft tree, so a wide, high-acceptance tree becomes memory-infeasible. We remove the snapshots. Using a tree-structured WY transform of the gated delta rule, we compute every draft node's output with a single triangular solve and reconstruct only the one accepted state on commit, storing a small pseudo-value matrix instead of per-node states; the derivation depends only on the gated delta rule, not on any other architectural detail. In serving benchmarks on two scales of one hybrid model family (Qwen3.5 35B and 397B) this cuts speculative recurrent-state memory and KV-cache pressure at identical acceptance length, turning the freed HBM into higher throughput and much lower time-to-first-token (TTFT) wherever memory binds, and costing a few percent where it does not. For tree width the same memory buys affordability: a wider, higher-acceptance draft becomes possible, though not yet a throughput win.
发表机构
- Thomson Reuters(汤森路透)
机构由 AI 辅助整理,请以论文原文为准。