发表机构
Tsinghua University; Southern University of Science and Technology; The University of Hong Kong; Peking University; Shenzhen University of Advanced Technology(清华大学; 南方科技大学; 香港大学; 北京大学; 深圳先进技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对循环Transformer解码昂贵问题,提出深度异步自推测(DAS),利用值成熟性异步读取前缀,实现4-7倍吞吐提升。
AI 中文摘要
循环Transformer在循环深度上重用共享块,使得自回归解码成本高昂,因为每个生成的令牌都需要多次顺序的循环传递。自推测解码器通过在早期深度草拟并在全深度验证来降低此成本,但通常将草拟计算绑定到来自相同循环深度的前缀表示。我们发现查询和键比值更早接近其最终深度表示,且受控的前缀通道干预表明,成熟值显著改善浅层草拟预测。受此不对称性启发,我们引入了深度异步自推测(DAS),它将草拟计算的深度与它读取的已验证前缀表示的深度解耦。其Mature-V原语允许浅层查询检索全深度前缀值,而无需额外的循环计算。我们进一步开发了DAS-Wave,它将深度异步前缀读取与携带的并行细化、渐进块增长和独立的全深度验证器相结合。在四个循环模型检查点以及数学和代码工作负载上,DAS-Wave在相同推理栈中相对于配对的全深度自回归解码实现了4.00--6.96倍的平均吞吐量提升。这些结果将前缀信息深度确定为循环自推测的有效设计轴。
英文摘要
Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce Depth-Asynchronous Self-Speculation (DAS), which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its Mature-V primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop DAS-Wave, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves 4.00--6.96$\times$ mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.