LoopSpec:循环Transformer的流水线式自推测解码
LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers
浏览论文内容
中文总结 AI 辅助
针对循环Transformer解码延迟高的问题,提出免训练的自推测解码框架LoopSpec,通过流水线重叠草稿生成与验证,并引入选择性第二提议,实现高达6.83倍推理加速。
中文摘要 AI 辅助
循环Transformer通过跨循环深度重复应用共享的Transformer块堆栈,以紧凑的参数规模实现了强大的性能。然而,与参数规模相当的标准Transformer模型相比,它们产生了更高的解码延迟,因为共享权重在每个循环深度都被访问。为了提高解码效率,自推测解码特别适合循环Transformer,因为它们的中间循环状态可以直接提供草稿预测,而无需辅助草稿模型。因此,我们提出了LoopSpec,一个为循环Transformer量身定制的免训练自推测解码框架。LoopSpec从早期循环状态中提取草稿令牌,并以流水线方式运行,将未来令牌的草稿生成与当前令牌的目标验证重叠。为了在不产生过多计算开销的情况下提高草稿准确性,我们引入了来自更深循环深度的选择性第二提议,同时确保在贪婪和采样两种机制下均实现无损解码。此外,我们以闭式形式推导了最优提议深度,并表明预测与测量结果一致。在推理和编码基准测试中,LoopSpec在多种循环Transformer上实现了高达6.83倍的推理加速。
英文摘要
Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur higher decoding latency than standard Transformer models of comparable parameter size because shared weights are accessed at every recurrent depth. To improve decoding efficiency, self-speculative decoding is particularly well suited to Looped Transformers, as their intermediate recurrent states can directly provide draft predictions without an auxiliary draft model. We therefore propose LoopSpec, a training-free self-speculative decoding framework tailored for Looped Transformers. LoopSpec extracts draft tokens from early recurrent states and operates in a pipelined manner, overlapping draft generation of future tokens with target verification of the current token. To improve draft accuracy without excessive compute overhead, we introduce a selective second proposal from deeper recurrent depth while ensuring lossless decoding under both greedy and sampling regimes. Furthermore, we derive the optimal proposal depths in closed form and show the prediction matches measurement. Across reasoning and coding benchmarks, LoopSpec achieves up to 6.83$\times$ inference speedup across diverse Looped Transformers.
发表机构
- Seoul National University(首尔大学)
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。