arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15533cs.LGcs.AI

机制可解释性的困境:一个形式化视角

The Misery of Mechanistic Interpretability: A Formal Perspective

Tobias Ladner, Matthias Althoff

首次发表
浏览论文内容

中文总结 AI 辅助

针对机制可解释性中IRN在对抗扰动下特征翻转的问题,提出形式化验证框架认证忠实性上界,并通过验证感知训练收紧界限,为LLM可解释性提供首个形式化保证。

中文摘要 AI 辅助

机制可解释性已成为理解前沿语言模型的主导视角,因为其内部运作复杂且本质上是黑箱。为了深入理解这些模型,可解释替代网络(IRN)在所有层上进行训练,通过稀疏激活的神经元暴露可解释的特征。然而,IRN的忠实性通常仅在干净数据上进行经验性评估,我们表明,即使语义上微小的输入扰动也会翻转主导的IRN特征——从而翻转人类可理解的理解——这一现象在五个开放权重模型家族(GPT-2 small、Gemma 2 2B、Gemma 3 1B、Llama 3.2 1B、R1-Distill-Qwen 1.5B)中均存在。我们提出了首个用于IRN忠实性的形式化验证框架,其中可达性分析在对抗性场景下认证了忠实性差距的可靠上界。此外,我们表明,对IRN进行验证感知训练可显著收紧这一认证界限,恢复了安全审计人员可以依据的特征级解释。综合来看,据我们所知,这些结果为大型语言模型的机制可解释性提供了首个形式化保证。

英文摘要

Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inherently black boxes. To gain insights into these models, interpretable replacement networks (IRNs) are trained at all layers, exposing interpretable features through sparsely activated neurons. However, the faithfulness of an IRN is usually evaluated only empirically on clean data, and we show that even semantically minor input perturbations flip the dominant IRN features-and thus the human-understandable interpretation-across five open-weight model families (GPT-2 small, Gemma 2 2B, Gemma 3 1B, Llama 3.2 1B, R1-Distill-Qwen 1.5B). We propose the first formal verification framework for the faithfulness of an IRN, where reachability analysis certifies a sound upper bound of the faithfulness gap in adversarial scenarios. Moreover, we show that verification-aware training of IRNs substantially tightens this certified bound, restoring a feature-level interpretation that safety auditors can act on. Together, these results give, to the best of our knowledge, the first formal guarantees for mechanistic interpretability of large language models.

发表机构

  • Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑