桥接路由头:多语言多跳推理在大型语言模型中的位置
Bridge Routing Heads: Where Multilingual Multi-hop Reasoning Lives in LLMs
浏览论文内容
中文总结 AI 辅助
本研究揭示多语言LLM中多跳推理依赖语言特异的桥接路由头,通过消融与放大实验证明激活干预可挽救跨语言推理失败。
中文摘要 AI 辅助
多语言大型语言模型(LLM)能用不同语言回答相同的多跳推理问题,但我们缺乏对其是否共享内部电路的机制性解释。我们通过一个三阶段流水线,在两个大型多语言LLM中识别出桥接路由头(Bridge Routing Heads, BRH)。所得的语言特定头集合在五种语言中表现出近乎完全的互斥性,Llama 3.1 70B的平均Jaccard相似度仅为0.017,Qwen 2.5 72B为0.057,揭示了语言特有的电路。消融通用BRH会使两跳负对数似然(NLL)比随机头基线增加39-89倍,提供了其作用的直接因果证据。在失败的目标语言推理中放大这些头,无需训练即可挽救高达51.7%的跨语言失败。这两个模型共享这种双电路模式,但头部分配不同:Llama将链式推理集中在一个大型通用池中,而Qwen则依赖更大的语言特定池。这些结果共同表明,仅通过激活层面的干预就能从跨语言推理失败中恢复正确答案。
英文摘要
Multilingual LLMs answer the same multi-hop reasoning question across languages, but we lack a mechanistic account of whether they share an internal circuit. We identify Bridge Routing Heads (BRH) in two large multilingual LLMs through a three-stage pipeline. The resulting language-specific head sets exhibit near-complete mutual exclusivity across the five languages, with a mean Jaccard similarity of only 0.017 for Llama 3.1 70B and 0.057 for Qwen 2.5 72B, revealing language-idiosyncratic circuits. Ablating general BRH increases two-hop Negative Log-Likelihood (NLL) by 39-89x the random-head baseline, providing direct causal evidence of their role. Amplifying these heads in a failing target-language pass rescues up to 51.7% of cross-lingual failures, with no training. The two models share this dual-circuit pattern but allocate heads differently: Llama concentrates chaining in a large general pool, while Qwen leans on larger language-specific pools. Together these results show that activation-level intervention alone can recover correct answers from cross-lingual reasoning failures.
发表机构
- Chosun University(朝鲜大学)
- Soongsil University(崇实大学)
机构由 AI 辅助整理,请以论文原文为准。