发表机构
Institute for Artificial Intelligence, Peking University; School of Mathematical Sciences, Peking University(北京大学人工智能研究院; 北京大学数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究循环Transformer中共享权重是否导致每循环执行相同操作,发现隐藏状态可路由共享计算至不同转移,注意力模式是因果通路。
AI 中文摘要
循环Transformer重复应用同一组Transformer层,从而形成一种用于潜在计算的循环架构。它们在迭代推理和长度泛化任务上的强劲表现暗示了一个有吸引力的解释:循环可能提供一种归纳偏置,使模型能够在各循环之间重用学到的算法。然而,权重共享本身并不意味着每个循环执行相同的操作。这引出一个基本问题:每个循环是否真的重复相同的计算?如果不是,是什么将共享参数路由到不同的操作?我们以图游走作为测试案例来研究这个问题。在模型的原生轨迹中,解码出的预测可以前进不同数量的图步,或者停留在已到达的目标处,这表明循环进展不必遵循固定的一循环一步模式。接着我们证明,通过修改进入的隐藏状态,可以引导一个冻结的循环转向不同的转移:一个学到的线性层$J$在不改变共享Transformer层的情况下选择期望的转移。为了测试这种引导机制的工作原理,我们使用激活修补,发现注意力模式可以恢复其效果并切换所选的转移。在五对匹配的图模型中,改变骨干训练期间的中间监督会改变$J$能够诱导的转移。这表明$J$选择的是骨干学到的计算,而非创造新的算法。综合这些结果,隐藏状态可以控制共享计算,而注意力路由是一个因果通路。
英文摘要
Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises a basic question: is each loop actually repeating the same computation, and if not, what routes the shared parameters to different operations? We study this question using graph walks as a test case. In the model's native trajectories, decoded predictions can advance by different numbers of graph steps or remain at a reached target, showing that recurrent progress need not follow a fixed one-loop-one-step pattern. We then show that a frozen loop can be steered toward different transitions by modifying its entering hidden state: a learned linear layer $J$ selects the desired transition without changing the shared Transformer layers. To test how this steering works, we use activation patching and find that attention patterns can recover its effects and switch the selected transition. Across five matched pairs of graph models, changing intermediate supervision during backbone training changes which transitions $J$ can induce. This suggests that $J$ selects computations learned by the backbone rather than creating new algorithms. Together, these results show that the hidden state can control shared computation, with attention routing as a causal pathway.
Comments36 pages. Code and reproduction materials: https://github.com/wjjpku/howloop