Circuit-Diff:基于事实编辑的干预方法用于定位归因图中的知识
Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs
浏览论文内容
中文总结 AI 辅助
本文提出Circuit-Diff方法,通过低秩事实编辑干预模型并识别归因图中变化特征,以定位与编辑知识相关的电路节点,并提供了因果验证和开源工具。
中文摘要 AI 辅助
机制可解释性将特征定义为神经网络的基本单元,将电路定义为执行其计算的加权子图。由于单个神经元具有多义性,跨层转码器(CLTs)被引入作为一种通过生成归因图来近似模型电路的方法。然而,该图的节点是未标记的特征:读取图意味着对其进行剪枝,然后手工确定每个存留节点的含义。为了使CLT更易于用于电路发现,我们引入了Circuit-Diff,它通过低秩事实编辑对模型本身进行干预,并将在该编辑下归因图中角色发生变化的特征视为与所编辑知识相关的特征。在我们检查的编辑中,被标记的节点不仅是对象令牌的检测器:从CLT发布的特征仪表板读取,它们包括与新旧对象相关的历史、地理和关联特征。我们形式化了该方法,测量了事实编辑后冻结的CLT保持可靠的程度,通过在多达24个CounterFact编辑上进行修补来因果测试所选节点,提供了一个案例研究,并发布了基于circuit-tracer包构建的开源实现,以及两个额外工具(多提示聚合和基于规则的超级节点标记)。
英文摘要
Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its computation. Because individual neurons are polysemantic, Cross-Layer Transcoders (CLTs) were introduced as a way to approximate a model's circuits by generating an attribution graph. The nodes of that graph, however, are unlabeled features: reading a graph means pruning it and then working out by hand what each surviving node means. To make CLTs easier to use for circuit discovery, we introduce Circuit-Diff, which intervenes on the model itself with a low-rank factual edit and takes the features whose role in the attribution graph changes under that edit as related to the edited knowledge. On the edits we examine, the flagged nodes are not only detectors of the object token: read off the CLT's released feature dashboards, they include features for the history, geography and associations surrounding the old and new objects. We formalize the method, measure how reliable a frozen CLT remains after a factual edit, test the selected nodes causally by patching them on up to 24 CounterFact edits, give a case study, and release an open-source implementation built on the circuit-tracer package, together with two further tools (multi-prompt aggregation and rule-based supernode labeling).
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。