图学习能否学习电路?
Can Graph Learning Learn Circuits?
浏览论文内容
中文总结 AI 辅助
该研究提出图电路学习(GCL)框架,将电路定位视为图机器学习问题,在扩充的InterpBench基准上评估GCL等方法,发现图机器学习为电路定位提供了新视角。
中文摘要 AI 辅助
电路定位是一种机械可解释性任务,目标是识别Transformer计算图中足以复现特定行为的稀疏子图。多数已建立的方法会针对每个模型-任务对独立定位电路,我们则将电路定位视为图机器学习问题,其中计算图的边代表计算路径,图神经网络(GNN)对这些路径间的交互进行建模。我们引入图电路学习(GCL),这是一种有监督的摊销框架,可在多个模型-任务对上训练GNN并将其应用于未见案例。为提供充足数据,我们用源自TracrBench程序的额外案例扩充InterpBench基准。在14种评估的GCL配置中,得分最高的在16个原始预留InterpBench案例上的中位数边AUROC为0.902(四分位区间[0.861, 0.942]),这接近已发表的InterpBench中EAP-IG的中位数0.910,但仍低于ACDC的0.959。移除所有消息传递边会使中位数降至0.825。我们还将GNN可解释性方法PGExplainer适配到电路定位,在相同案例上获得中位数边AUROC为0.858。这些初步结果表明,图机器学习为电路定位提供了自然且潜在强大的视角,我们希望这一视角能促进两个领域间更紧密的交流。
英文摘要
Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient to reproduce a particular behavior. Most established methods localize circuits independently for each model--task pair. We instead frame circuit localization as a graph machine learning problem in which the edges of a computation graph represent computational pathways, and graph neural networks (GNNs) model interactions among these pathways. We introduce Graph Circuit Learning (GCL), a supervised, amortized framework that trains a GNN across multiple model--task pairs and applies it to unseen cases. To provide sufficient data, we augment the InterpBench benchmark with additional cases derived from the TracrBench programs. Of the 14 evaluated GCL configurations, the highest scored a median edge AUROC of $0.902$ (interquartile interval $[0.861, 0.942]$) on the 16 original held-out InterpBench cases. This is close to the published InterpBench median of $0.910$ for EAP-IG while remaining below ACDC's $0.959$. Removing all message-passing edges reduces the median to $0.825$. We also adapt PGExplainer, a GNN explainability method, to circuit localization, obtaining a median edge AUROC of $0.858$ on the same cases. These preliminary results suggest that graph machine learning offers a natural and potentially powerful perspective on circuit localization, and we hope this perspective encourages closer exchange between the two communities.