内在结构:机制可解释性的谱可识别性
Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability
浏览论文内容
中文总结 AI 辅助
本研究提出机制可解释性的谱可识别性定理,利用Koopman算子构建模型内在指纹,在GPT-2 small等模型上验证了谱的收敛性及相关性质,为机制可解释性的基础组件提供了首个可识别性理论。
中文摘要 AI 辅助
机制可解释性通过识别模型内部的电路来解释模型,但无法区分电路是模型的固有属性还是发现该电路的方法的人工产物。稀疏自编码器说明了这一问题:不同的随机种子和宽度会从相同的激活中恢复出差异显著的特征,且没有理论表明这种变异性是偶然的还是结构性的。我们将用于可解释性的字典学习建立在可识别性的基础上,将前向传播视为以深度为时间的受控动力系统,并通过Koopman算子对其进行提升,得到一个有限线性实现,其谱是模型的与坐标无关的属性。我们证明,在排列不变的情况下,该谱可从M个校准样本中以M^{-1/2}的速率恢复——据我们所知,这是首个针对机制可解释性基础组件的可识别性定理,还包含匹配的极小极大下界、针对重尾激活的中位数均值变体,以及一个解离定理:当该实现非正规时,承载激活方差的方向与跨深度承载信息的方向无法重合。可识别对象与清晰可解释对象并非同一事物。在GPT-2 small、Gemma-2-2B和Qwen3-8B-Base上,该谱在所有位置均收敛,且在Qwen3-8B-Base上达到了预测指数(0.506±0.031);不足部分会针对每个单元的样本阈值汇聚成一条曲线。Koopman模式在间接对象识别上优于随机方向,但落后于主成分,且差距随深度距离衰减4.1倍,与定理预测一致。Koopman谱是具有明确误差棒的可识别、模型内在指纹,而非清晰可解释的分解。
英文摘要
Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variability is incidental or structural. We put dictionary learning for interpretability on an identifiability footing. Treating the forward pass as a controlled dynamical system with depth as time and lifting it with the Koopman operator yields a finite linear realisation whose \emph{spectrum} is a coordinate-free property of the model. We prove the spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation - to our knowledge the first identifiability theorem for a mechanistic-interpretability primitive, with a matching minimax lower bound, a median-of-means variant for heavy-tailed activations, and a dissociation theorem: whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide. The identifiable object and the legible object are not the same object. On GPT-2 small, Gemma-2-2B and Qwen3-8B-Base the spectrum converges everywhere and attains the predicted exponent on Qwen3-8B-Base ($0.506 \pm 0.031$); shortfalls collapse onto one curve against each cell's sample threshold. Koopman modes beat random directions but lose to principal components on indirect-object identification, with the gap decaying $4.1\times$ in depth-distance, as the theorem predicts. The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition.
发表机构
- IISER Bhopal(印度博帕尔科学教育与研究所)
- IBM Research(IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。