发表机构
University of Southern California(南加利福尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CircuitSteer是利用SAEs识别多层语义回路的新型框架,通过几何对齐合成引导向量实现多点干预,在多任务多模型上,其引导效果优于牺牲流畅性或覆盖不足的对比方法。
AI 中文摘要
控制大型语言模型(LLMs)的行为仍是AI对齐领域的关键挑战。现有引导方法如对比激活加法(CAA)通常依赖从聚合激活差异中推导的固定单层干预,这类方法对语义多样的输入施加单一干预,往往无法在各层维持一致的行为变化,限制了引导效果。本研究提出CircuitSteer,这一利用稀疏自编码器(SAEs)识别和操纵分布在多层的连贯语义回路的新型框架。通过构建基于特征共激活和解码器方向几何对齐的特征流回路,我们分离出负责目标行为的特定多层子回路,随后从这些稀疏特征中合成密集引导向量,并应用多点干预以引导模型的内部语义轨迹。我们在涵盖毒性、情绪强度、奉承、弃权(不执行)等任务的多样化对比示例上,对CircuitSteer进行了评估,涉及两类模型家族。在所有模型和数据集上,CircuitSteer是唯一能始终保持流畅性的干预方法;对比方法要么牺牲文本质量,要么覆盖不足,在奉承、弃权(不执行)等复杂行为上完全失效。这些结果表明,通过对所选特征施加几何对齐实现的多层回路引导,比静态单点干预能产生更鲁棒、更有效的行为控制。代码可在该https URL获取。
英文摘要
Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions derived from aggregate activation differences. These methods impose a single intervention across semantically diverse inputs and often fail to sustain consistent behavioral changes across layers, limiting the effectiveness of the steering. In this work, we introduce CircuitSteer, a novel framework that leverages Sparse Autoencoders (SAEs) to identify and manipulate coherent semantic circuits distributed across multiple layers. By constructing a feature flow circuit based on feature co-activation and the geometric alignment of decoder directions, we isolate the specific multi-layer subcircuits responsible for a target behavior. We then synthesize dense steering vectors from these sparse features and apply multi-point interventions to guide the model's internal semantic trajectory. We evaluate CircuitSteer using contrastive examples across a diverse set of tasks, including toxicity, emotion-intensity, sycophancy, and refusal, spanning two model families. Across all models and datasets, CircuitSteer is the only method to consistently produce fluency-preserving interventions; competing methods either sacrifice text quality or lack coverage, failing entirely on complex behaviors like sycophancy and refusal. These results demonstrate that multi-layer circuit steering, enabled by enforcing geometric alignment among selected features, yields strictly more robust and effective behavioral control than static single-point interventions. Code is available at https://github.com/mehrshad-sdtn/CircuitSteer.