arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

S^3martCirc:自监督智能电路发现

S^3martCirc: Self-supervised Smart Circuit Discovery

Wendy Zheng, Yinhan He, Liang Wu, Jundong Li

arXiv 2609.00755首次发表:更新:

发表机构

University of Virginia; Nokia(弗吉尼亚大学; 诺基亚)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大型语言模型机制可解释性的两阶段范式缺陷,提出S^3martCirc统一框架,联合发现电路与功能角色,实验显示其电路发现性能优于现有方法。

AI 中文摘要

大型语言模型(LLMs)在文本摘要、问答等多种任务中展现出卓越性能,但其黑箱特性掩盖了内部决策过程。机制可解释性(MI)旨在通过将神经网络逆向工程为人类可理解的算法解决这一问题。当前针对LLMs的MI方法通常遵循两阶段范式:第一阶段识别重要组件(电路发现),组件通常是注意力头或前馈神经元等单个节点;第二阶段确定其在特定任务中的作用(功能解释)。然而,这种顺序方法忽略了一个基本见解:组件的重要性与其功能角色本质上相互依赖。统一这些阶段面临两个关键挑战:(1)功能角色通常与特定节点或组件绑定,限制了泛化性;(2)其识别依赖主观解释而非可量化指标。为应对这些挑战,我们提出S^3martCirc(自监督智能电路发现,Self-supervised Smart Circuit Discovery),这是一个同时进行电路发现与功能解释的统一框架。S^3martCirc将节点行为抽象为跨任务泛化的两类通用计算角色,并定义了分配这些角色的量化指标,使重要性与功能角色可联合发现而非按顺序进行。大量实验表明,我们的框架在电路发现任务中优于现有方法。

英文摘要

Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks, from text summarization to question answering. Despite these capabilities, their black-box nature obscures internal decision-making processes. Mechanistic interpretability (MI) aims to address this by reverse-engineering neural networks into human-understandable algorithms. Current MI approaches for LLMs typically follow a two-stage paradigm: first identifying important components (circuit discovery), where components are typically individual nodes such as an attention head or feedforward neuron, and second determining the role they play in a certain task (functional interpretation). However, this sequential approach overlooks a fundamental insight: a component's importance and its functional role are inherently codependent. Unifying these stages presents two key challenges: (1) functional roles are often tied to specific nodes or components, limiting generalization, and (2) their identification relies on subjective interpretation rather than quantifiable metrics. To address these challenges, we propose S^3martCirc (Self-supervised Smart Circuit Discovery), a unified framework that simultaneously discovers circuits and interprets functionality. S^3martCirc abstracts node behavior into two general computational roles that generalize across tasks and defines a quantitative metric for assigning them, enabling importance and functional role to be discovered jointly rather than in sequence. Extensive experiments show that our framework outperforms existing methods in circuit discovery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑