AI 中文总结
该研究提出机制层析方法,通过设计测量恢复模型内部机制与干预效应,在不同模型访问场景下验证其有效性,还在GPT-2-small IOI、Qwen-2.5-7B上完成相关实验。
AI 中文摘要
机制可解释性旨在探寻模型未直接暴露的量:表征状态、组件效应、交互作用以及对干预的响应。补丁法(Patching)、梯度、赫森向量积(Hessian-vector products)和子集干预在不同的访问假设下提供不同的测量,且可能针对不同的量。我们将它们的共享测量结构形式化为机制层析:用于恢复内部机制和干预效应的设计测量。对于选定的基和干预族,测量形式为 y = Ax + w,其中 A 描述干预,x 是目标映射,w 包含非线性响应、采样误差和基的误设定。该方法给出了实用流程:从成本最低的测量开始,在预期规模的保留干预上测试,校准简单的不匹配,当存在结构化残差时扩展测量族。控制提供了严苛的验证设置,因为指导干预的估计值充当观察者。在双隐马尔可夫模型(two-HMM model)中,控制误差随观察者误差上升,而目标改进可隐藏干扰状态的移动。在仅前向访问下,稀疏聚合测量比坐标补丁法用更少的干预恢复有限效应映射;在梯度访问下,有限探针改进局部归因映射;提升测量和赫森向量积恢复一阶映射遗漏的交互,而 Tracr 表明所需的族依赖于基。在 GPT-2-small IOI 上,Name Mover-负 Name Mover 交互是三个测试跨组对中最大的保留预测项;在 Qwen-2.5-7B 上,有限校准使加性拒绝响应映射足够,故保留误差不支持成对提升。
英文摘要
Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities. We formulate their shared measurement structure as mechanistic tomography: designed measurement for recovering internal mechanisms and intervention effects. For a chosen basis and intervention family, measurements take the form y = Ax + w, where A describes the interventions, x is the target map, and w contains nonlinear response, sampling error, and basis misspecification. This language gives a practical procedure: start with the least costly measurements, test on held-out interventions at the intended scale, calibrate simple mismatch, and expand the measurement family when structured residuals remain. Control provides a demanding validation setting because an estimate that guides an intervention acts as an observer. In a two-HMM model, control error rises with observer error, while target improvement can hide nuisance-state movement. Under forward-only access, sparse aggregate measurements recover a finite-effect map with fewer interventions than coordinate patching. With gradient access, finite probes improve a local attribution map. Lifted measurements and Hessian-vector products recover interactions missed by first-order maps, while Tracr shows that the required family depends on the basis. On GPT-2-small IOI, the Name Mover-Negative Name Mover interaction is the largest held-out predictive term among three tested cross-group pairs. On Qwen-2.5-7B, finite calibration makes an additive refusal-response map adequate, so held-out error does not support pairwise lifting.
Comments24 pages, 13 figures