干预梯度与检测激活:语言模型中的目标特征学习
Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出目标特征学习框架,比较六种方法(基于激活值、梯度及不同估计器)与SAEs在15个任务上的表现,发现激活值方法检测最优,梯度方法干预最优,二者互补。
中文摘要 AI 辅助
模型内部特征可以通过其识别特定概念的能力以及在被操纵(例如通过引导或权重编辑)时的因果效应来研究。特征学习的一种主要方法是稀疏自编码器(SAEs),它学习广泛的特征字典,这些字典与特定概念的关系通常事后识别。然而,许多可解释性问题反而是假设驱动的,并且关注预先指定的概念。我们将这种设置研究为目标特征学习,即为这种预定义概念构建单个特征。我们提出了对三种模型信号(激活值、激活梯度和参数梯度)和两种估计器(对比均值和学习的单维编码器-解码器)的受控比较,产生了六种目标方法,其中CAA和GRADIEND是现有实例,四种新方法覆盖其余组合。我们将这些方法与预训练的SAEs在15个任务和三种语言模型上进行比较,评估检测和因果干预。跨模型来看,最强的检测性能由对比激活值方法实现,而最强的干预性能由基于梯度的方法实现。总体而言,我们的结果表明,目标特征质量共同取决于模型信号和估计器,检测和干预捕捉互补的属性。
英文摘要
Model-internal features can be studied through both their ability to identify a specified concept and their causal effect when manipulated, e.g., through steering or weight editing. A prominent approach to feature learning is Sparse Autoencoders (SAEs), which learn broad feature dictionaries whose relation to particular concepts is typically identified post hoc. However, many interpretability questions are instead hypothesis-driven and concern a concept specified in advance. We study this setting as targeted feature learning, where a single feature is constructed for such a predefined concept. We present a controlled comparison across three model signals (activation values, activation gradients, and parameter gradients) and two estimators (contrastive mean and a learned one-dimensional encoder-decoder), yielding six targeted methods, with CAA and GRADIEND as existing instances and four new methods covering the remaining combinations. We compare these methods against pretrained SAEs across 15 tasks and three language models, evaluating both detection and causal intervention. Across models, the strongest detection performance is achieved by contrastive activation value methods, whereas the strongest intervention performance is achieved by gradient-based methods. Overall, our results show that targeted feature quality depends jointly on the model signal and estimator, with detection and intervention capturing complementary properties.