发表机构
China Financial Certification Authority (CFCA); PetroChina Southwest Oil and Gasfield Company(中国金融认证中心; 中国石油西南油气田公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过植入隐藏通路使因果干预的有害分歧可检验,发现现有距离指标无法识别,提出HPC测试(AUROC≥0.99),并揭示优化干预会主动利用隐藏通路。
AI 中文摘要
因果干预(如激活修补和分布式对齐搜索(DAS))是对神经网络进行机制性断言的主要工具。近期研究表明,这些干预通常会推动表征偏离模型的自然分布,而这种分歧有时无害,有时则有害:它可能招募模型在自然输入上从未使用的通路,从而通过错误的机制产生预期答案。目前尚无方法能区分这两种情况。我们通过在预训练语言模型中植入隐藏通路,使这一问题变得可检验;这些通路在构造上对所有基准提示保持静默,因此哪些干预依赖于它们可以被精确获知。在GPT-2 small上的72种配置和100,800次干预中,我们发现三点:(i)干预位点的最近邻和局部PCA距离(如先前工作所用)在识别通过植入通路给出正确答案的干预时,得分低于随机水平(AUROC 0.35-0.47)。(ii)隐藏通路贡献(HPC)——一种无标签测试,将下游单元钳制到具有相同输出的自然运行状态,并测量决策消失的程度——在通路表现为单元级越界活动时,能以AUROC >= 0.99标记通路主导的干预,但当每个单元都保持在其自然范围内时则失败,我们将此确定为开放问题。(iii)优化后的干预主动寻找隐藏通路:在性别任务上,DAS在四个家族中的三个中,将其90-95%的成功通过植入通路实现,而下游流形内惩罚可将此比例降至5%以下,代价是成功率下降6-11个百分点。在未修改的GPT-2中,成功的干预几乎不表现出单元级越界依赖。
英文摘要
Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural distribution, and that such divergence is sometimes harmless and sometimes pernicious: it can recruit pathways the model never uses on natural inputs, so that an intervention produces the expected answer through the wrong mechanism. No method currently tells the two cases apart. We make this question testable by planting hidden pathways inside pretrained language models; the pathways are silent on every benchmark prompt by construction, so which interventions depend on them is known exactly. Across 72 configurations and 100,800 interventions on GPT-2 small, we find three things. (i) Nearest-neighbour and local-PCA distances at the intervention site, as used in prior work, score below chance (AUROC 0.35-0.47) at picking out interventions that give the right answer through a planted pathway. (ii) Hidden-Pathway Contribution (HPC), a label-free test that clamps downstream units to the regime of natural runs with the same output and measures how much of the decision disappears, flags pathway-dominated interventions with AUROC >= 0.99 when the pathway shows up as unit-level out-of-regime activity, but fails when every unit stays within its natural range, which we identify as the open problem. (iii) Optimised interventions actively seek hidden pathways: on a gender task, DAS routes 90-95% of its successes through planted pathways for three of four families, and a downstream on-manifold penalty cuts this share to under 5% at a cost of 6-11 points of success rate. In unmodified GPT-2, successful interventions show almost no unit-level out-of-regime reliance.
Comments13 pages, 1 figure, 6 tables. Beiming Liu and Minjie Chen contributed equally