发表机构
College of Computing and Data Science, Nanyang Technological University, Singapore; Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore; School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学计算机与数据科学学院; 南洋理工大学跨学科研究生项目; 南洋理工大学电气与电子工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出条件共消融(CoAx)方法,通过测量在移除主要组件后剩余单元消融效应的增长,来发现变压器电路中因自修复而隐藏的备份组件,在GPT-2-small IOI电路上将备份头恢复的ROC-AUC从0.33提升至0.91。
AI 中文摘要
机制可解释性通常依赖于组件级干预来发现模型如何产生行为。这指导了下游的归因、能力消除和模型剪枝,通过单独消融每个单元的效果进行评分。当组件重要性是加性时,这种一阶评分是自然的,但当变压器自修复时,它会变得具有误导性:在主要组件被移除后,一个休眠的备份可以接管,掩盖主要组件的测量效果,而备份本身在完整模型中看起来无关。我们将这种失败重新定义为恢复任务,即条件电路补全,并引入条件共消融(CoAx),这是一种无标签、输出基础的评分,询问在移除主要组件集后,每个剩余单元的消融效果增长了多少。这种条件增长揭示了一阶评分丢弃的二阶交互。在GPT-2-small IOI电路上,CoAx将备份头恢复的ROC-AUC从0.33提高到0.91,优于所有基线,包括自修复感知的梯度评分(最佳0.82);反事实修补验证了恢复的头因果地承载了修复。相同的无标签程序迁移到八个模型上的归纳。除了发现之外,恢复的备份纠正了自修复掩盖的归因,识别了能力消除所需的组件,并产生了从124M到7B的修复感知结构化剪枝。因此,组件重要性不仅仅是孤立单元的性质:在鲁棒电路中,重要的组件只有在使其必要的干预下才能变得可见。
英文摘要
Mechanistic interpretability seeks to explain transformer behavior through circuits: sets of internal components that causally support a behavior. However, self-repair creates a blind spot: ablating a primary component can activate a dormant backup, so a circuit that explains behavior in the intact model can become incomplete under the intervention used to test it. We formulate this gap as conditional circuit completion: given a primary set, identify components that become causally important after its removal. We introduce conditional co-ablation (CoAx), which ranks candidates by growth in ablation effect after primary-set removal. We show that a perfectly dormant backup can be indistinguishable from an irrelevant component to per-unit intact-state scores, whereas its conditional effect change exactly aggregates all interaction orders linking it to the removed set. On GPT-2-small's Indirect Object Identification (IOI) circuit, CoAx recovers the documented backup heads at 0.941 ROC-AUC, versus 0.815 for the strongest intact-state attribution baseline and 0.758 for the matched conditional-energy control. Recovery drops to 0.40 +/- 0.13 AUC for alternative component sets matched in behavioral effect, output displacement, and depth, showing that recovery is specific to the removed circuit. Beyond recovery, the CoAx-selected heads are causally load-bearing: freezing them after primary removal sharply reduces the IOI margin, while adding them to the incomplete circuit reduces incompleteness from 0.75 to 0.21. More broadly, conditional growth aligns with intervention-derived repair in 11/12 held-out instances across 4 mechanism clusters, and CoAx completions outperform matched random completions on all 8 non-GPT-2 models spanning 6 architecture families. Together, causal explanations of self-repairing transformers must account for backup circuitry when primary components fail.