循环Transformer中的递归推理调度
Scheduling Recursive Reasoning in Looped Transformers
- Southern University of Science and Technology(南方科技大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出轨迹自适应进展-波动调度器(TAPS),通过动态调整循环Transformer中递归更新的步长,平衡持续进展与波动,从而在无需重训练的情况下提升推理准确率并实现最高1.56倍加速。
AI中文摘要:
递归推理模型因扩展测试时计算而日益受到关注,通常通过共享参数迭代细化潜在状态。然而,这些模型以固定的单位尺度应用每次学习到的更新,当更新持续取得进展时可能过于保守,而在更新波动时则可能过于激进,从而限制了额外循环的收益。为了理解尺度应如何沿轨迹变化,我们首先分析了终端损失对递归更新尺度的敏感性。我们证明其时间平均值可精确分解为持续进展和中心化波动两部分贡献。基于此,我们提出了轨迹自适应进展-波动调度器(TAPS),该调度器跟踪递归更新中两者的平衡并在线调整步长。理论上,我们建立了TAPS在减少期望终端损失并以更少递归循环达到目标质量方面的充分条件。实验上,我们表明TAPS无需重新训练即可提升结构化推理任务的终端准确率。通过进一步将进展-波动原则融入训练,TAPS在匹配基线准确率的情况下额外获得准确率提升,并实现高达1.56倍的墙钟加速。TAPS的广泛适用性得到其在多种递归架构和推理策略中有效性的支持。综合来看,这些结果确立了更新尺度作为递归推理中与架构和深度互补的控制轴。
英文摘要:
Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss to recurrent update scale. We show that its temporal average admits an exact decomposition into persistent-progress and centered-fluctuation contributions. Based on this, we introduce the Trajectory Adaptive Progress-Fluctuation Scheduler (TAPS), which tracks their balance across recurrent updates and adapts the step size online. Theoretically, we establish sufficient conditions under which TAPS reduces expected terminal loss and reaches a target quality in fewer recurrent loops. Empirically, we show that TAPS improves terminal accuracy across structured reasoning tasks without retraining. By further incorporating the progress-fluctuation principle into training, TAPS yields additional accuracy gains with up to 1.56 times wall-clock speedup at matched baseline accuracy. The broad applicability of TAPS is supported by its effectiveness across diverse recurrent architectures and inference strategies. Together, these results establish update scale as complementary control axis of recurrent inference alongside architecture and depth.