发表机构
University of Toronto; McGill University(多伦多大学; 麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示机械可解释性中目标级恢复差距:干预定义的忠实度可能偏好行为复现较差的电路,导致发现算法错误排序,恢复信号可修复多数错误排序。
AI 中文摘要
机械可解释性旨在恢复负责模型行为的内部计算。自动电路发现的进展通常被框定为搜索问题:更好的归因或优化应能识别出更好的机制。这假设评估目标能在找到电路后识别出更好的电路。我们表明,基于干预定义的忠实度反而可能偏好一个同等规模但复现模型行为较差的电路,从而产生目标级恢复差距。在四个人类参考任务和InterpBench上,我们在固定普通重采样下,将验证忠实度与保留提示上的行为进行比较。行为标准是与完整模型(包括其错误)的一致性,但在Greater-Than任务上,我们使用语义准确性。受控参考编辑揭示了无需任何发现算法的错误排序,而EAP、EAP-IG、ACDC和Edge-SP的输出表现出相同的失败。在重采样下,KL在这些方法上对人类参考任务中9.4%-41.2%的候选对进行了错误排序。我们研究上下文失真作为解释:替换被排除的信号会改变保留组件所操作的输入。从接收者的完整模型执行中恢复所选信号,修复了发现池中100个持久KL错误排序中的96个,在验证和保留提示上均如此。电路及其原始行为分数保持不变。这些发现表明,当目标奖励错误的候选时,仅靠更好的发现是不够的。
英文摘要
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
Comments34 pages, 2 figures