大型语言模型从强化学习中学到了什么?基于固定SAE追踪的机制可解释性视角
What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track
- Peking University(北京大学)
- National University of Singapore(新加坡国立大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出固定SAE追踪框架,发现RL主要增强现有特征而非创造新特征,且引导这些特征可恢复约80%的性能提升。
AI中文摘要:
强化学习(RL)在大型语言模型训练中被广泛用于提升特定能力,然而RL如何重塑模型仍鲜为人知。先前解释RL工作机制的尝试大多提供行为层面的视角,而未揭示RL在表征层面给模型带来了什么:RL能否创造真正新颖的特征,以及它增强或抑制了哪些现有特征?机制可解释性的最新进展表明,稀疏自编码器(SAEs)是分解内部激活为人类可解释特征的有力工具;然而,它们无法直接应用于追踪训练过程中的变化。在本工作中,我们提出了固定SAE追踪(Fixed-SAE Track)框架,该框架在基础模型和所有RL检查点的激活池上,为每个考虑的层训练一个共享的SAE,保持所有特征方向固定,从而通过可解释的SAE潜在变量的激活严格定义表征偏移,包括检测新出现的特征。在多个数据集和RL算法上的验证表明,RL引起的偏移是微小、渐进、概念特定且集中在后期层,主要提高了少量阶梯令牌的采样率,以及诸如步骤分隔符和答案定界符之类的格式脚手架,而非重塑问题内容。将这些特征引导到基础模型中可恢复RL约80%的性能提升,表明RL主要激发模型已具备的能力,正如引导所做的那样。我们进一步设计了一个特征构造已知的合成基准,以测试RL能否灌输真正新颖的特征。我们相信固定SAE追踪为追踪表征偏移提供了一种原则性方法,并为理解强化学习如何改变LLM的内部表征提供了表征层面的证据。
英文摘要:
Reinforcement learning (RL) is widely utilized in large language model training to improve targeted capabilities, yet how RL reshapes a model remains poorly understood. Prior attempts to explain how RL works largely offer behavioral perspectives, leaving open what RL gives a model at the representation level: can RL create genuinely novel features, and which existing features does it enhance or suppress? Recent developments in mechanistic interpretability suggest sparse autoencoders (SAEs) as a promising lens to decompose internal activations into human-interpretable features; however, they cannot be directly applied to tracking change across training. In this work, we introduce Fixed-SAE Track, a framework that trains one shared SAE per considered layer on activations pooled across the base model and all RL checkpoints, holding every feature direction fixed so that representation shifts are rigorously defined through the activations of interpretable SAE latents, including the detection of emerging novel features. Validated across multiple datasets and RL algorithms, we find that RL-induced drift is small, gradual, concept specific, and concentrated in late layers, mainly enhancing the sampling rates of a small set of ladder tokens, formatting scaffolding such as step breaks and answer delimiters, rather than reshaping problem content. Steering these features into the base model recovers around 80% of RL's performance gain, suggesting that RL primarily elicits capabilities the model already possesses, much as steering does. We further design a synthetic benchmark with features known by construction to test whether RL can instill genuinely novel features. We believe Fixed-SAE Track provides a principled approach to tracking representation shifts and offers representational evidence for understanding how reinforcement learning changes the inner representation of LLMs.