ScopeSAE:具有可解释层选择的模型范围特征发现
ScopeSAE: Model-Scope Feature Discovery with Interpretable Layer Selection
浏览论文内容
中文总结 AI 辅助
本文提出ScopeSAE,通过归一化梯度归因为每个token选择预测相关子空间来训练稀疏自编码器,实现重建优于原始的效果,并提升特征利用率、可解释性和字典冗余度。
中文摘要 AI 辅助
稀疏自编码器(SAEs)是机制可解释性中的核心工具。然而,现有的SAEs主要按层进行训练,其建模子空间因此由层身份固定,独立于哪些token-层状态实际驱动每个预测。我们认为,这一约束导致了逐层SAEs中观察到的若干局限性,包括特征利用率低、字典冗余度高,以及特征缺乏直接的行为依据。在本文中,我们提出了ScopeSAE,它通过基于归一化梯度的归因方法,将每个预测归因于其最具影响力的token-层状态,从而为每个token选择建模子空间,并在由此产生的与预测相关的子空间上学习特征。实验上,ScopeSAE产生了一种我们称之为“重建优于原始”的效果。SAE的回写重建产生的下一token交叉熵低于原始激活,据我们所知,这一结果此前尚未在SAEs中报道过。通过干预分析和KL微调对照实验,我们表明这种效果可归因于ScopeSAE的预测相关子空间本身,而非架构变化。ScopeSAE进一步在有效特征数量、可解释性、利用率和字典冗余度方面优于现有的基于层的基线,这表明通过预测相关性选择SAE建模子空间能够产生更有用且行为上有意义的特征。
英文摘要
Sparse autoencoders (SAEs) are a central tool in mechanistic interpretability. However, existing SAEs are primarily trained per layer. The modeling subspace is therefore fixed by layer identity, independent of which token-layer states actually drive each prediction. We argue that this constraint contributes to several limitations observed in layer-wise SAEs, including low feature utilization, high dictionary redundancy, and features that lack direct behavioral grounding. In this paper, we propose ScopeSAE, which selects the modeling subspace per token by attributing each prediction to its most influential token-layer state via normalized gradient-based attribution, and learns features over the resulting prediction-relevant subspace. Empirically, ScopeSAE yields an effect we term reconstruction-better-than-original. Written-back reconstructions of the SAE produce lower next-token cross-entropy than the original activations, an outcome that, to our knowledge, has not previously been reported for SAEs. Through interventional analyses and a KL fine-tuning counter-experiment, we show that this effect is attributable to ScopeSAE's prediction-relevant subspace itself rather than to architectural changes. ScopeSAE further improves effective feature count, interpretability, utilization, and dictionary redundancy over existing layer-based baselines, suggesting that choosing the SAE modeling subspace by predictive relevance leads to more useful and behaviorally meaningful features.
发表机构
- The University of Sydney(悉尼大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- DeepWisdom(深度智慧)
- University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。