发表机构
University of Virginia; Nokia(弗吉尼亚大学; 诺基亚)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型语言模型机制可解释性的两阶段范式缺陷,提出S^3martCirc统一框架,联合发现电路与功能角色,实验显示其电路发现性能优于现有方法。
AI 中文摘要
大型语言模型(LLMs)在文本摘要、问答等多种任务中展现出卓越性能,但其黑箱特性掩盖了内部决策过程。机制可解释性(MI)旨在通过将神经网络逆向工程为人类可理解的算法解决这一问题。当前针对LLMs的MI方法通常遵循两阶段范式:第一阶段识别重要组件(电路发现),组件通常是注意力头或前馈神经元等单个节点;第二阶段确定其在特定任务中的作用(功能解释)。然而,这种顺序方法忽略了一个基本见解:组件的重要性与其功能角色本质上相互依赖。统一这些阶段面临两个关键挑战:(1)功能角色通常与特定节点或组件绑定,限制了泛化性;(2)其识别依赖主观解释而非可量化指标。为应对这些挑战,我们提出S^3martCirc(自监督智能电路发现,Self-supervised Smart Circuit Discovery),这是一个同时进行电路发现与功能解释的统一框架。S^3martCirc将节点行为抽象为跨任务泛化的两类通用计算角色,并定义了分配这些角色的量化指标,使重要性与功能角色可联合发现而非按顺序进行。大量实验表明,我们的框架在电路发现任务中优于现有方法。
英文摘要
Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks, from text summarization to question answering. Despite these capabilities, their black-box nature obscures internal decision-making processes. Mechanistic interpretability (MI) aims to address this by reverse-engineering neural networks into human-understandable algorithms. Current MI approaches for LLMs typically follow a two-stage paradigm: first identifying important components (circuit discovery), where components are typically individual nodes such as an attention head or feedforward neuron, and second determining the role they play in a certain task (functional interpretation). However, this sequential approach overlooks a fundamental insight: a component's importance and its functional role are inherently codependent. Unifying these stages presents two key challenges: (1) functional roles are often tied to specific nodes or components, limiting generalization, and (2) their identification relies on subjective interpretation rather than quantifiable metrics. To address these challenges, we propose S^3martCirc (Self-supervised Smart Circuit Discovery), a unified framework that simultaneously discovers circuits and interprets functionality. S^3martCirc abstracts node behavior into two general computational roles that generalize across tasks and defines a quantitative metric for assigning them, enabling importance and functional role to be discovered jointly rather than in sequence. Extensive experiments show that our framework outperforms existing methods in circuit discovery.