发表机构
Basis Research Institute; University of Massachusetts Amherst; MIT(基础研究所; 马萨诸塞大学阿默斯特分校; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出WISE估计量和JuntaLearner方法,用于在语言模型中高效发现考虑组件间交互的电路,显著提升电路识别性能并支持大规模模型。
AI 中文摘要
将行为定位到语言模型的单个组件是机制可解释性的核心目标。然而,一次只对一个组件进行评分会忽略上下文相关效应:一个主要组件可能抑制备用组件的激活,从而导致组件排序问题。实际因果关系通过见证者来研究此类交互的结构:见证者是提供上下文信息以解决交互项的变量。然而,使用见证者的估计通常需要组合枚举,在实践中不可行。我们引入了见证者集成集效应(WISE),这是一族因果估计量,建立在见证者机制之上,同时对原因和见证者的集合取期望,以保持计算可行性。基于这种方法,我们提出了JuntaLearner,一种基于梯度的电路发现方法,它学习根据组件在不同大小的组件和见证者集合中的因果影响对组件进行排序。除了保真度指标外,我们还引入了必要性和任务特异性度量,以及电路识别分数(CRS),用于汇总各电路大小的每个指标,同时强调由小电路实现的效果。在任务和规模递增的模型中,JuntaLearner在所有指标上相比归因基线获得了更高的平均CRS。由于其成本不随候选组件数量增长,JuntaLearner可以扩展到大型模型,同时考虑集合级交互并避免一阶近似。
英文摘要
Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies the structure of such interactions via witnesses: variables that provide contextual information to resolve interaction terms. However, estimation with witnesses typically requires combinatorial enumeration and is infeasible in practice. We introduce the witness-integrated set effect (WISE), a family of causal estimands that build on the witness mechanism while taking expectations over sets of causes and witnesses to remain computationally feasible. Building on this approach, we introduce JuntaLearner, a gradient-based circuit discovery method that learns to rank components by their causal impact across varying-sized sets of components and witnesses. Alongside faithfulness metrics, we introduce measures of necessity and task specificity, and the circuit recognition score (CRS) to summarize each metric across circuit sizes while emphasizing effects achieved by small circuits. Across tasks and models of increasing size, JuntaLearner achieves higher mean CRS compared to attribution baselines on all metrics. Since its cost does not grow with the number of candidate components, JuntaLearner scales to large models while accounting for set-level interactions and avoiding first-order approximations.