发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多跳推理中非循环Transformer的深度局部存储问题,提出DiscoLoop架构,通过循环同时携带离散嵌入和连续隐藏状态,实现近乎完美的泛化,并在符号和合成语言任务中显著减少训练步数。
AI 中文摘要
大型语言模型在允许将中间步骤外部化为思维链(CoT)时,在许多推理任务上取得了强劲性能。然而,许多问题要求模型在生成答案之前,在单次前向传递中内化多步推理。我们通过两跳推理研究这一挑战,这是一个代表性任务,模型必须在单次前向传递中组合多条参数化知识。标准的非循环Transformer存在深度局部存储问题:早期层学习的事实对于第二跳检索发生的位置不可用。我们发现循环Transformer通过重用相同的内存缓解了这个问题,但泛化仍然不完美。我们表明剩余的瓶颈是表示性的。在两跳推理任务中,第一个循环通常使正确的桥接实体几乎完美可解码,但相应的隐藏状态与桥接令牌嵌入仍然对齐不良。令人惊讶的是,一种简单的无训练重对齐干预几乎消除了泛化差距。基于这一见解,我们提出了DiscoLoop,一种循环架构,其循环同时携带离散嵌入通道和连续隐藏状态通道。DiscoLoop在符号和合成语言多跳推理任务上以显著更少的训练步数实现了近乎完美的准确性。当应用于实际预训练时,DiscoLoop比循环Transformer基线获得了更低的训练损失和更强的基准性能,表明混合通道设计可迁移到实际语言建模。
英文摘要
Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). However, many questions require the model to internalize the multi-step reasoning within a single forward pass before generating the answer. We study this challenge through two-hop reasoning, a representative task where the model must compose multiple pieces of parametric knowledge within a single forward pass. Standard non-recurrent Transformers suffer from a depth-local storage problem: facts learned in earlier layers are unavailable where second-hop retrieval happens. We found that Looped Transformers mitigate this issue by reusing the same memory, but still generalize imperfectly. We show that the remaining bottleneck is representational. In the two-hop reasoning task, the first loop often makes the correct bridge entity nearly perfectly decodable, yet the corresponding hidden state remains poorly aligned with the bridge token embedding. Surprisingly, an easy training-free realignment intervention nearly closes the generalization gap. Building upon this insight, we propose DiscoLoop, a looping architecture whose recurrence carries both a discrete embedding channel and a continuous hidden-state channel. DiscoLoop achieves near-perfect accuracy with substantially fewer training steps across symbolic and synthetic-language multi-hop reasoning tasks. When applied to real-world pretraining, DiscoLoop attains lower training loss and stronger benchmark performance than looped-transformer baselines, suggesting that the mixed-channel design transfers to practical language modeling.
Comments16 pages, 7 figures