发表机构
Kyoto University(京都大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对激活转向中复合方向纠缠问题,提出转向向量解剖框架,通过训练稀疏自编码器解缠转向向量,实现精确控制模型行为。
AI 中文摘要
激活转向已成为一种轻量级的推理时方法,用于控制大型语言模型(LLM)的行为。然而,用于干预LLM激活的传统转向向量,例如通过均值差法导出的那些向量,往往会将多个语义和风格概念纠缠到一个复合方向中,导致不可预测的转向效果。我们的核心目标是将这一复合方向解缠为其组成概念。为此,我们提出了转向向量解剖(Steering Vector Dissection),一个框架,用于从这些复合方向中显式地分离出单个且语义一致的特征。具体来说,我们将正负激活配对并取其差值,以生成一组实例级转向向量,并直接在它们上训练一个专门的稀疏自编码器(SAE)。在两个数据集、两个模型和两种干预深度上的定量评估表明,我们的方法产生了一组语义一致的基向量,其转向效果彼此可区分。此外,我们展示了这种解缠能够实现对模型行为的精确控制。
英文摘要
Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional steering vectors used to intervene in LLMs' activations, such as those derived from the difference-in-means method, tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects. Our core objective is to disentangle this composite direction into its constituent concepts. To this end, we propose Steering Vector Dissection, a framework to explicitly isolate individual and semantically consistent features from these composite directions. Specifically, we pair positive and negative activations and take their differences to generate a set of instance-level steering vectors, and train a dedicated Sparse Autoencoder (SAE) directly on them. Quantitative evaluations across two datasets, two models, and two intervention depths show that our method yields a set of semantically consistent basis vectors whose steering effects are mutually distinguishable. Furthermore, we show that this disentanglement enables precise control over model behaviors.
Comments31 pages, 5 figures