arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM 层立即相互纠正

LLM Layers Immediately Correct Each Other

Arjun Patrawala, Jiahai Feng, Erik Jones, Jacob Steinhardt

arXiv 2609.07876首次发表:更新:

发表机构

University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现相邻 Transformer 层通过 TLCM 机制系统性相互纠正,该机制在预训练中出现并自适应调节,可解释 SAE 特征低特异性及转码器优势。

AI 中文摘要

近期语言模型可解释性研究中的方法采用诸如稀疏自编码器等技术,将残差流的贡献分解为线性、语义上有意义的特征。这类方法通常被解释为识别残差流中持续存在且后续层在其基础上构建的特征。我们通过识别 Transformer 层纠正机制(TLCM)来挑战这一观点,在该机制中,相邻的 Transformer 层系统地抵消彼此贡献的某些部分。TLCM 出现在 7 个主要开源模型家族中的 5 个中,并在多样文本中几乎所有 token 上被激活。我们表明,TLCM 在预训练期间出现,对上下文相关的 token 作用最强,并根据前一层的输出自适应地校准其纠正强度。利用层雅可比矩阵,我们进一步表明,TLCM 选择性地纠正特定子空间,同时增强其他子空间,我们通过一个“提出-拒绝”框架对此进行解释,在该框架中,层提出候选特征,后续层选择性地移除不合适的特征。这种动态表明,任何层的残差流都包含瞬态提议以及持续特征,这有助于解释为什么 SAE 特征描述通常特异性较低,为什么有效的模型引导需要极端的特征放大,以及为什么转码器在理论上优于 SAE。

英文摘要

Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linear, semantically meaningful features. Such methods are commonly interpreted as identifying features that persist in the residual stream and that subsequent layers build upon. We challenge this view by identifying the Transformer Layer Correction Mechanism (TLCM), wherein adjacent transformer layers systematically counteract portions of each other's contributions. TLCM appears in 5 out of 7 major open-source model families and activates across nearly all tokens in diverse texts. We show that TLCM emerges during pretraining, operates most strongly on contextually dependent tokens, and adaptively calibrates its correction strength based on the preceding layer's output. Using the layer Jacobian, we further show that TLCM selectively corrects specific subspaces while reinforcing others, which we interpret through a ``propose-and-reject'' framework in which layers propose candidate features and subsequent layers selectively remove inappropriate ones. This dynamic suggests that the residual stream at any layer contains transient proposals alongside persistent features, helping explain why SAE feature descriptions often have low specificity, why effective model steering requires extreme feature amplification, and why transcoders hold a theoretical advantage over SAEs.

CommentsPublished at NeurIPS 2025

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑