arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27656cs.LGcs.CL

以源为中心的状态演化循环Transformer

Looped Transformers with Source-Centered State Evolution

Bum Jun Kim, Kohei Hayashi, Shunsuke Kamiya, Masanori Koyama, Yusuke Iwasawa, Yutaka Matsuo

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出以源为中心的状态演化(SCSE)循环Transformer,解决输入条件与共享循环参考的协调问题,在多个文本任务中提升了可控循环质量,验证了其设计的有效性。

中文摘要 AI 辅助

循环Transformer通过在循环深度上重复使用同一个Transformer块,在参数数量固定的情况下增加了有效深度,从而创建了有用的训练和测试时计算轴。然而,该共享块必须控制在训练深度和外推深度上变化的隐藏状态的整个轨迹。此外,在加法注入式循环Transformer中,输入条件信号会在每个循环步骤重新引入,因此在输入条件参考下应用共享变换仍可移动隐藏状态。本文提出以源为中心的状态演化(SCSE),旨在协调输入条件与保持参考的共享循环。具体而言,SCSE通过其学习到的锚点和初始偏差保留输入依赖性,允许非零偏差驱动循环计算,同时将零偏差映射为零,并通过其零偏差掩码保证精确的锚点不变性。由此,指定的锚点在构造上是单步不动点。零偏差强制偏置是从锚点自身产生的下一个偏差,在SCSE中消失,而非零偏差保持活跃并支持依赖状态的循环计算。我们的理论表明,零偏差强制偏置是一种设计自由度,其任务效应可能有害、中性或有益;SCSE通过将偏置设为零来选择精确的锚点不变性,从而解决了这一选择。在WikiText-2、WikiText-103、直接网络语料预训练、保留网络文本迁移及LAMBADA补全任务中,SCSE提升了可控循环质量的前沿水平。消融研究表明,学习到的锚点和锚点坐标偏差循环是性能提升的主要贡献因素,训练后的模型案例研究将锚点响应诊断建立在观测到的循环运动基础上。

英文摘要

Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective depth at a fixed parameter count. However, that shared block must then govern an entire trajectory of varying hidden states over trained and extrapolated depths. Furthermore, in additive-injection looped Transformers, an input-conditioned signal is reintroduced at every recurrent step, so applying the shared transition at an input-conditioned reference can still move the hidden state. In this paper, we propose Source-Centered State Evolution (SCSE), which is designed to reconcile input conditioning with reference-preserving shared recurrence. Specifically, SCSE retains input dependence through its learned anchor and initial deviation, allows nonzero deviations to drive recurrent computation while mapping zero deviation to zero, and guarantees exact anchor invariance through its zero-deviation mask. The designated anchor is thereby a one-step fixed point by construction. The zero-deviation forcing bias is the next deviation produced from the anchor itself and vanishes in SCSE, while nonzero deviations remain active and support state-dependent recurrent computation. Our theory shows that the zero-deviation forcing bias is a design degree of freedom whose task effect can be harmful, neutral, or beneficial; SCSE resolves this choice in favor of exact anchor invariance by setting the bias to zero. Across WikiText-2, WikiText-103, direct web-corpus pretraining, held-out web-text transfer, and LAMBADA completion, SCSE improves the controlled recurrent quality frontier. Ablation studies identify the learned anchor and the anchor-coordinate deviation recurrence as the primary contributors to the gain, and a trained-model case study grounds the anchor-response diagnostic in observed recurrent motion.

补充信息

↑