发表机构
Xi’an Jiaotong University; University of Illinois Urbana-Champaign(西安交通大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对语言模型两跳查询失败问题,在受控环境训练Transformer,揭示其泛化规律与跨层不匹配机制,提出循环式训练策略以提升分布外两跳泛化能力。
AI 中文摘要
大型语言模型(LLMs)可解决复杂多跳问题,但在简单两跳查询上却出现令人困惑的失败:尽管模型可能正确存储每一跳的信息,却常无法将二者结合。为理解该现象的内部机制,我们在受控符号环境中从头训练Transformer。实验揭示两跳泛化的规律:当第二跳符合训练分布时,模型泛化可靠;但当第二跳偏离分布时,模型总会失败。通过机制分析,我们为这些不同的泛化行为提供完整解释:在模型泛化成功的场景中,性能由跨上下文同一实体的一致中间表示的出现所驱动;而在第二跳分布外的场景中,失败源于跨层不匹配:底层正确构建这些中间表示,但上层虽在对应原子事实上接受训练,却主要学习将它们映射到输出,而非对其进行推理。基于此见解,我们提出一种循环式训练策略,使Transformer能在不同输入形式间复用其推理电路,大幅提升分布外两跳查询的泛化能力。
英文摘要
Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctly store each individual hop, it often fails to combine them. To understand the internal mechanisms of this phenomenon, we train transformers from scratch in a controlled symbolic environment. Our experiments reveal a pattern in two-hop generalization: models generalize reliably when the second hop follows the training distribution, but always fail when it deviates. Through mechanistic analysis, we provide a complete explanation for these distinct generalization behaviors: in settings where models generalize successfully, performance is driven by the emergence of consistent intermediate representations for the same entities across contexts, whereas failures on settings where the second hop is out-of-distribution arise from a mismatch across layers: lower layers correctly construct these intermediate representations, but upper layers, while trained on corresponding atomic facts, primarily learn to map them to outputs rather than to reason over them. Driven by this insight, we propose a recurrent-style training strategy, which enables transformers to reuse their reasoning circuitry across input forms and substantially improves generalization on out-of-distribution two-hop queries. Our data and code are available at https://github.com/zzl-strong/two_hop .
CommentsEMNLP 2026, findings