发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LoopCD,一种无需训练的对比解码框架,利用循环Transformer的中间状态作为引导信号,在减少推理计算的同时提升解码质量,并支持循环次数减半。
AI 中文摘要
循环Transformer通过在循环中重复执行共享模块来实现参数效率。每个循环产生一个可解码为相同下一个令牌的中间表示,但标准解码会丢弃早期状态。由于早期循环包含较少的计算,循环性固有地提供了对齐的弱-强预测对,无需辅助模型或外部训练。我们引入了LoopCD,一种无需训练的对比解码框架,通过将最终预测与早期循环传递进行对比来指导令牌选择,可在逻辑空间中进行一次额外输出传递(LoopCD-Logits),或在隐藏状态空间中零输出开销(LoopCD-Hidden)。在四个循环Transformer家族中,LoopCD在完整循环深度下提供了显著且一致的改进:LoopCD-Logits将Ouro-2.6B-Thinking的AIME 2024 pass@1从61.88%提升至73.33%,而LoopCD-Hidden将Huginn的HumanEval pass@1从22.56%提升至31.71%。至关重要的是,这些性能提升使得循环次数减半成为可能,同时仍能匹配或超过完整深度的无引导基线,将前向FLOPs减少22.5%至48.2%。通过将中间循环状态转化为有效的引导信号,LoopCD在显著减少推理计算的同时实现了更优的解码质量。
英文摘要
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.
Comments32 pages, 19 figures