arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23033cs.LG

波前解码:循环语言模型的并行自推测解码

WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Hyeongju Ha, Jae-Joon Kim

AI总结:

针对循环语言模型解码延迟高的问题,提出波前解码(WFD),利用权重共享和混合深度状态并行化草拟与验证,在多个基准上实现显著加速。

AI中文摘要:

循环语言模型通过重复应用一个权重共享的块来增加有效深度,而无需增加参数数量,但每个生成令牌所需的T次顺序循环块调用会显著增加解码延迟。为解决此问题,我们引入了波前解码(WFD),这是一种为循环语言模型设计的免训练自推测解码框架。WFD利用了这些架构的两个特性:中间循环输出提供了有效的草稿预测,且权重共享允许不同位置和循环深度的令牌状态在一次批处理的循环块调用中处理。WFD将这些混合深度状态组织成对角波前,在浅层持续草拟新位置,同时将较早位置推进至全深度验证。与相位分离的草拟-然后-验证调度不同,WFD因此在相同的循环调用中共同批处理草拟和验证,而被拒绝的草稿则使用全深度预测进行纠正。在六个Spec-Bench任务类别中,WFD在Ouro-2.6B上实现了2.42倍加速,在Huginn-3.5B上实现了3.54倍加速,相较于自回归解码,始终优于草拟-然后-验证方法。跨循环KV共享进一步减少了波前KV流量,并将WFD在Huginn-3.5B上的加速提升至4.81倍。

英文摘要:

Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore concurrently batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B. The code is available at https://github.com/summerbro-hhj/wavefront-decoding.

补充信息

↑