发表机构
University of Minnesota(明尼苏达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较了五种GPTNeoX变体在多跳推理任务上的表现,发现潜在推理模型通过稀疏循环搜索电路实现更好的深度泛化,优于思维链等方法。
AI 中文摘要
大型语言模型可以通过不同形式的中间计算(从基于标记的轨迹到在潜在空间中执行的计算)执行多步推理并提高任务性能。然而,一个问题仍然悬而未决:这些不同形式的思考是否依赖于相同的底层机制?为了解决这个问题,我们在一个扩展的多跳推理任务(ProsQA-Ext)上从头训练并比较了同一个GPTNeoX骨干网络的五个变体:一个普通模型、一个思维链(CoT)模型、一个暂停标记模型,以及两个端到端优化且没有中间推理轨迹的潜在推理模型。我们发现,强大的分布内(ID)性能并不能保证深度泛化。普通模型、思维链模型和暂停标记模型能很好地解决ID问题,但主要依赖局部图特征,并且对跳数更长的分布外(OOD)问题泛化能力较差。相比之下,潜在变体泛化更好,并表现出与图上前向可达性传播一致的内部动态。因果干预和电路分析将该计算定位到瓶颈潜在模型中的一个稀疏循环搜索电路:一个注意力头检索图关系,一个MLP和残差流在循环步骤中更新可达性状态,而多个注意力头则共同执行候选匹配。总之,这些结果表明,不同的思考机制可以学习到不同的计算解决方案,即使在相似的ID性能下也是如此。在这种设置中,潜在循环支持一种可重用的前向搜索算法,该算法能够泛化到训练深度之外。
英文摘要
Large Language Models can perform multi-step reasoning and improve task performance through different forms of intermediate computation, from token-based traces to computation carried out in latent space. However, a question remains open: do these different forms of thinking rely on the same underlying mechanism? To address this, we train and compare five variants of the same GPTNeoX backbone from scratch on an extended multi-hop reasoning task (ProsQA-Ext): a vanilla model, a Chain-of-Thought (CoT) model, a Pause Token model, and two latent-reasoning models that are optimized end-to-end without intermediate reasoning traces. We find that, strong in-distribution (ID) performance does not guarantee depth generalization. Vanilla, CoT, and Pause Token models solve ID problems well, but rely largely on local graph features and generalize poorly to out-of-distribution (OOD) problems with longer hops. In contrast, latent variants generalize better and show internal dynamics consistent with forward reachability propagation on the graph. Causal interventions and circuit analysis localize this computation to a sparse recurrent search circuit in the bottleneck latent model: an attention head retrieves graph relations, an MLP and the residual stream update the reachability state across recurrent steps, while multiple attention heads together then do the candidate matching. Together, these results show that different thinking mechanisms can learn distinct computational solutions, even at similar ID performance. In this setting, latent recurrence supports a reusable forward-search algorithm that generalizes beyond the training depth.