arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30954cs.AI

FTB图:多语言语言模型中首词广播器与语言身份头部电路的确定与验证

FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models

Arjun Pillai, Christian Hoang, Anjelo Laroza

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过EAP与精确激活修补,在六种多语言模型上提取首词语言身份电路,发现深层广播枢纽及预训练主导的路由机制,并指出梯度近似需精确验证。

中文摘要 AI 辅助

在多语言环境中运行的大型语言模型必须在生成早期解决目标响应语言的问题,然而控制首词语言身份决策的因果电路仍未被充分映射。我们提出了一种端到端的结构电路分析,涵盖四个家族的六种模型架构:GPT-2、BLOOM-560M、Pythia-1B/2.8B 以及 Qwen2.5-1.5B Base/Instruct。使用带有 FP16 主动钳位的边归因修补(EAP),随后通过精确激活修补验证,并设置 2,000 个候选边的搜索上限,我们提取了驱动首词语言广播的有向无环图。在独立模型中,我们观察到深层或中深层广播枢纽,但证据在 Pythia-2.8B 和 BLOOM-560M 上最强,因为 GPT-2 和 Pythia-1B 留下较少的图外头部用于比较,而两个 Qwen2.5-1.5B 变体则反转了必要性检查。从 Pythia-1B 扩展到 2.8B 增加了节点参与度,同时保持相似的已验证边预算,产生更稀疏的拓扑结构。Qwen2.5-1.5B 基础版和指令版的电路保留了 84.7% 的 Jaccard 相似度,包括第 27 层枢纽,表明首词路由主要在预训练期间建立,并由指令微调保留。最后,EAP 分数与大多数模型的精确修补差异相关性较弱,表明在 FP16 中线性梯度近似可能偏离因果干预,这促使精确修补验证成为可靠电路发现的必要步骤。

英文摘要

Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four families: GPT-2, BLOOM-560M, Pythia-1B/2.8B, and Qwen2.5-1.5B Base/Instruct. Using Edge Attribution Patching (EAP) with FP16 active clamping, followed by exact activation patching verification with a 2,000-candidate-edge search ceiling, we extract directed acyclic graphs driving first-token language broadcasting. Across the standalone models, we observe deep or mid-to-deep broadcasting hubs, though the evidence is strongest for Pythia-2.8B and BLOOM-560M because GPT-2 and Pythia-1B leave few out-of-graph heads for comparison, while both Qwen2.5-1.5B variants invert the necessity check. Scaling from Pythia-1B to 2.8B expands node participation while maintaining a similar verified edge budget, producing sparser topology. The Qwen2.5-1.5B base and instruct circuits retain 84.7% Jaccard similarity, including the Layer 27 hub, indicating that first-token routing is largely established during pretraining and preserved by instruction tuning. Finally, EAP scores correlate weakly with exact patching deltas across most models, showing that linear gradient approximations can diverge from causal interventions in FP16 and motivating exact-patching verification for reliable circuit discovery.

发表机构

  • Irvington High School(欧文顿高中)
  • Gen4AIE

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑