AI 中文总结
本研究通过机制分析揭示循环Transformer长度泛化的两种机制及其失效原因,并提出循环状态编码计算状态、可转移残差方向控制信息参与后续计算的共同原则,表明长度泛化无需忠实逐步推理。
AI 中文摘要
循环Transformer能够泛化到比训练时遇到的更长的推理链,但实现这一行为的计算及其限制范围仍不清楚。我们从机制上比较了两种循环Transformer配置,分别称为匹配循环循环Transformer(MR-Loop)和解耦循环循环Transformer(DR-Loop),这反映了它们各自的循环训练方案。我们使用详细的机制分析评估了多项式迭代、有限状态组合和知识图谱遍历。注意力分析、中间状态解码和因果干预揭示了在最终答案监督下学习到的不同机制。MR-Loop在固定读出点更新中间状态,同时通过相邻令牌交互和可转移的进度线索推进关系选择。DR-Loop则跨关系位置传播中间状态,形成前进的计算前沿。然而,两种机制在更深层次上都变得不可靠:MR-Loop表现出其读出状态和进度线索的退化,而DR-Loop表现出状态传播可靠性下降。有限的自校正允许局部错误持续存在并累积。在两种模型中,我们发现了一个共同的表示原则:循环状态不仅编码任务相关内容,还编码其计算状态,即该内容是否仍以能够支持后续计算的形式存在。可转移的实时消耗和新鲜老化残差方向因果控制所表示的信息是否能参与后续计算,包括在训练视界之外。我们进一步表明,长度泛化不必依赖于忠实的逐步推理,因为循环Transformer可以利用任务结构而无需显式表示每个中间状态。
英文摘要
Looped Transformers can generalize to reasoning chains longer than those encountered during training, but the computations enabling this behavior and limiting its extent remain unclear. We mechanistically compare two looped-Transformer configurations, which we call the Matched-Recurrence Looped Transformer (MR-Loop) and Decoupled-Recurrence Looped Transformer (DR-Loop), reflecting their respective recurrence-training schemes. We evaluate polynomial iteration, finite-state composition, and knowledge-graph traversal using detailed mechanistic analysis. Attention analysis, intermediate-state decoding, and causal interventions reveal distinct mechanisms learned under final-answer supervision. MR-Loop updates an intermediate state at a fixed readout while advancing relation selection through adjacent-token interactions and a transferable progress cue. DR-Loop instead propagates intermediate states across relation positions, forming an advancing computational frontier. However, both mechanisms become unreliable at greater depths: MR-Loop exhibits degradation of its readout state and progress cues, while DR-Loop exhibits declining reliability of state propagation. Limited self-correction allows local errors to persist and compound. Across both models, we uncover a common representational principle: recurrent states encode not only task-relevant content but also its computational status, whether that content remains in a form that can support subsequent computation. Transferable live-consumed and fresh-aged residual directions causally control whether represented information can participate in subsequent computation, including beyond the training horizon. We further show that length generalization need not rely on faithful step-by-step reasoning, as Looped Transformers can exploit task structure without explicitly representing every intermediate state.
Comments39 pages, 10 figures