发表机构
Hasso Plattner Institute; IT University of Copenhagen(哈索·普拉特纳研究所; 哥本哈根信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现SAE特征流发现中解码器余弦相似度在MLP更新场景下失效,通过构建转移图谱和消融验证,在Pythia和Gemma模型中揭示了大量低余弦相似度但强效应的转移,为表示-更新机制诊断提供新工具。
AI 中文摘要
基础模型越来越多地通过微调、模型编辑和对齐流程进行适配,同时保留先前获得的能力。因此,理解支持这些适配的内部计算对于模型的持续演化变得越来越重要。稀疏自编码器(SAE)为残差流激活和子层输出提供了可解释的特征字典,但状态特征和更新特征如何相互作用以产生下游残差特征仍不清楚。在这项工作中,我们以MLP更新作为第一个测试案例。我们构建了一个三元组 $s_k + u_j \rightarrow t_\ell$ 的转移图谱,其中残差状态特征和MLP更新特征共同预测目标残差特征,并通过消融解码后的更新特征来验证候选三元组。在20M-token的Pythia-160M $L_7 \rightarrow L_8$ 运行中,我们发现了38,125个强消融效应转移,但其中88.0%的状态-目标和更新-目标解码器余弦相似度均低于0.7。作为初步的跨模型检查,20M-token的Gemma-3-4B $L_{21} \rightarrow L_{22}$ 运行通过消融解码后的更新特征,仅对排名前30,000的候选三元组进行了因果验证,其中53.6%的强效应三元组的状态-目标和更新-目标解码器余弦相似度均低于0.7。Gemma的结果在方向上与Pythia一致,但较弱,因为更新-目标余弦相似度恢复了许多最强的Gemma效应,且该运行并非完整的图谱因果验证。最终,我们的结果表明,特征流图谱可以作为表示-更新机制的诊断工具,从而为引导模型更新的工具提供信息。未来的工作将验证跨层、跨模型和跨SAE家族的更复杂模式。
英文摘要
Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities. Understanding the internal computations that support these adaptations is therefore becoming increasingly important for continual model evolution. Sparse autoencoders (SAEs) provide interpretable feature dictionaries for residual-stream activations and sublayer outputs, but it remains unclear how state features and update features interact to produce downstream residual features. In this work, we focus on MLP updates as a first test case. We construct a transition atlas of triples $s_k + u_j \rightarrow t_\ell$, where a residual-state feature and an MLP-update feature jointly predict a target residual feature, and validate candidate triples by ablating the decoded update feature. In a 20M-token Pythia-160M $L_7 \rightarrow L_8$ run, we find 38,125 strong ablation-effect transitions, but 88.0% have both state-target and update-target decoder cosine similarity below 0.7. As a preliminary cross-model check, a run of 20M-token Gemma-3-4B $L_{21} \rightarrow L_{22}$ causally validates only the top 30,000 ranked candidate triples by ablating the decoded update feature, and 53.6% of strong-effect triples have both state-target and update-target decoder cosine similarity below 0.7. The Gemma result is directionally consistent with Pythia, but weaker, since update-target cosine recovers many of the strongest Gemma effects and the run is not a full-atlas causal validation. Ultimately, our results suggest that feature flow atlases can serve as diagnostics of representation-update mechanisms and thereby inform tools for steering model updates. Future work will validate more complex patterns across layers, models, and SAE families.
Comments4 pages, 2 figures