发表机构
University of Toronto; McGill University; Tsinghua University(多伦多大学; 麦吉尔大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究质疑电路解释的有效性,发现现有电路能复现模型成功却遗漏大部分错误,提出以精确错误复现作为评估电路解释的必要条件。
AI 中文摘要
机制可解释性(MI)旨在通过分析模型的内部计算来解释其行为;基于电路的解释旨在通过紧凑的子网络来隔离这些计算,并通过消融模型的其余部分进行验证。我们表明,以这种方式验证的电路可能无法恢复模型行为的基本机制,因为它们能紧密复现模型的成功决策,却无法解释其大部分错误。此类解释应同时考虑模型的特定错误及其成功。我们通过在IOI、Docstring以及机制可解释性基准中的六个模型-任务设置上,分别测量模型成功与失败时的精确答案一致性(跨电路大小和消融设置)来评估这一要求。我们发现,许多测试的电路能紧密复现正确行为,却遗漏了模型的大部分错误。在GPT-2 small的间接宾语识别(IOI)任务中,在均值消融下,手工电路和测试的自动电路(包括一个针对模型完整输出分布训练的电路)在模型正确回答的提示上与模型的一致性为97.3%-99.5%,但在错误上仅为11.4%-41.7%。一个IOI案例研究表明,通过恢复被省略的注意力头可以找回丢失的错误,这些注意力头将错误复现率从14.2%提高到75.1%(在单独的留出集上),同时正确一致性仅下降0.41个百分点,超过了匹配的随机扩展和标量偏置控制。干预轨迹显示,被省略的计算如何为可复现的错误子集产生特定的错误答案。总之,这些发现表明电路可以在不充分解释模型失败的情况下保持任务成功,并支持将精确错误复现作为基于电路解释模型行为的必要但非充分测试。
英文摘要
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.