电路浓缩:将行为因果电路集中化的训练后方法
Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit
浏览论文内容
中文总结 AI 辅助
本文提出Circuit Condensation方法,通过训练后处理将模型行为的因果电路集中,在4种行为8个模型的实验中大幅缩减电路规模,且能保留任务性能,可用于提升机械可解释性的电路分析效率。
中文摘要 AI 辅助
机械可解释性的一种方法是通过电路(承载行为的组件与连接)来解释模型行为。现有冻结发现方法通常会返回数百条边,导致这些边难以检查、比较或全面验证。本文提出了电路浓缩(Circuit Condensation),一种将模型训练后处理以将行为集中到更小因果图中的方法。每一轮会修剪低归因边,并训练低秩适配器(low-rank adapter),通过剩余边匹配原始模型,仅在任务性能和通用能力保持不变时保留修剪操作。在4种行为和8个模型的实验中,32个设置里有30个的浓缩电路比最强的冻结基线更小,平均缩小8.1倍,最大达316倍。若不进行权重更新仅重复搜索,32个设置里有29个会产生更大的电路,表明是权重更新而非单纯搜索驱动了缩减。对19条电路的所有子集测试发现,11条无法再缩减,其余存在可移除边;成对消融实验揭示了边间的依赖关系,说明其效果无法独立理解。在间接宾语识别任务中,电路浓缩隔离出24个注意力头,其中17个有已记录的作用,而匹配的冻结电路对应61个注意力头、36个无记录作用;该浓缩电路是已发表机制的充分子电路而非其重构,能追踪原始模型的下一个token分布并预测其错误。
英文摘要
One approach to mechanistic interpretability explains behavior through circuits: the components and connections that carry it. Frozen discovery often returns hundreds of edges, making them hard to inspect, compare, or verify exhaustively. We introduce Circuit Condensation, which post-trains models to concentrate behaviors into smaller causal graphs. Each round prunes low-attribution edges and trains a low-rank adapter to match the original through what remains, retaining the cut only if task performance and general capability survive. Across four behaviors and eight models, condensed circuits are smaller than the strongest frozen baseline in 30 of 32 settings, by $8.1\times$ on average and up to $316\times$. Repeating the search without weight updates produces larger circuits in 29 of 32 settings, showing that weight updates, rather than search alone, drive the reduction. Testing every subset of 19 circuits finds 11 that cannot be reduced and reveals removable edges in the rest. Pair ablations expose dependencies between edges, showing that their effects cannot be understood independently. On indirect object identification, condensation isolates 24 heads, 17 of them with documented roles, against 61 heads and 36 undocumented ones for the matched frozen circuit: a sufficient sub-circuit of the published mechanism rather than a reconstruction of it. The resulting circuit tracks the original model's next-token distribution and predicts its errors.
发表机构
- George Mason University(乔治梅森大学)
机构由 AI 辅助整理,请以论文原文为准。