循环Transformer作为优化器
Looped Transformers as Optimizers
浏览论文内容
中文总结 AI 辅助
本文提出将循环Transformer的隐藏状态视为快速权重,从优化视角推导循环转换,并设计OperLoop方法,在匹配计算量下提升生成与常识推理性能。
中文摘要 AI 辅助
循环Transformer通过重复应用共享的Transformer块,提供了一种参数高效的深度扩展方法。最近的推理模型同样强调了通过更长的计算轨迹来扩展测试时计算的价值。然而,设计有效循环转换的原则仍未被充分理解。我们将循环隐藏状态视为在整个深度中更新的快速权重。我们将循环转换公式化为基于局部梯度的更新,其中循环块在每个深度预测隐式目标。我们的框架从投影、局部目标和优化器更新规则中以闭式推导出循环转换。将代表性的循环转换映射到该框架中,揭示了其转换与投影之间的不匹配。我们首先对齐现有转换的输入映射。然后,我们推导出OperLoop,它结合了显式权重衰减、自适应步长和delta目标。对齐的变体降低了训练损失并提高了平均常识准确率。在匹配的训练FLOPs下,OperLoop在平均生成性能上优于所比较的循环和非循环基线。这些结果支持了该框架对循环设计的实用性。我们将分析扩展到其他循环模型,并概述了未来循环转换设计的路线图。
英文摘要
Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longer computation trajectories. However, the principles for designing effective loop transitions remain poorly understood. We view the looped hidden state as a fast weight that is updated throughout the depth. We formulate loop transitions as local gradient-based updates, with recurrent blocks predicting implicit targets at each depth. Our framework derives loop transitions in closed form from a projection, a local objective and an optimizer update rule. Mapping representative loop transitions into this framework reveals mismatches between their transitions and projections. We first align the input maps of existing transitions. We then derive OperLoop, which combines explicit weight decay, adaptive step size and a delta objective. The aligned variants reduce training loss and improve average commonsense accuracy. OperLoop improves average generative performance over the compared looped and non-looped baselines under matched training FLOPs. These results support the framework's usefulness for loop design. We extend the analysis to additional loop models and outline a roadmap for future loop transition design.
发表机构
- HKUST (GZ)(香港科技大学(广州))
- SJTU(上海交通大学)
- UCAS(中国科学院大学)
- THU(清华大学)
- StepFun(阶跃星辰)
机构由 AI 辅助整理,请以论文原文为准。