arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25270cs.CLcs.AIcs.LG

转向信号来自何处:激活转向中的激活源选择

Where Steering Signals Come From: Activation Source Selection in Activation Steering

Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan

首次发表
浏览论文内容

中文总结 AI 辅助

研究激活转向中转向信号的来源选择,通过实验表明改变源激活会影响转向成功率,有效转向信号来自执行边界状态,基于此提出尾部减法,揭示转向取决于模型即将做的事而非已出现的内容。

中文摘要 AI 辅助

激活转向通过在推理时向隐藏状态添加向量或特征来控制语言模型,但这些转向信号的上游来源通常被视为次要细节。我们将此来源选择作为激活源选择进行研究,即源上下文和激活读出策略的组合,用于收集构建转向信号的隐藏状态。在保持下游干预不变的情况下,我们在三个指令调整模型和四个转向任务族中表明,仅改变源激活会显著改变转向成功率。我们还发现,有效的转向不能简单地由所需行为是否出现在源文本中来解释。相反,强信号来自执行边界状态,即模型即将产生或继续目标行为的状态。这种实现前/后的区别解释了基于答案的源有时有效的原因:它们的有用部分与执行边界方向对齐,而不仅仅是目标外观。基于此观点,我们引入了尾部减法,它从边界状态中去除共享的提示和延续语义,产生更清晰、更稳定的转向信号。总体而言,我们的结果表明,转向取决于模型即将做什么的表示,而不仅仅取决于已经出现的内容。

英文摘要

Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation source selection: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built. Holding the downstream intervention fixed, we show across three instruction-tuned models and four steering task families that changing only the source activations substantially changes steering success. We further find that effective steering is not explained simply by whether the desired behavior appears in the source text. Instead, strong signals come from execution-boundary states, where the model is about to produce or continue the target behavior. This pre-/post-realization distinction explains why answer-based sources sometimes work: their useful component aligns with execution-boundary directions rather than target appearance alone. Building on this view, we introduce tail subtraction, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals. Overall, our results suggest that steering depends on representations of what the model is about to do, not merely on what has already appeared.

发表机构

  • Peking University(北京大学)
  • Tsinghua University(清华大学)
  • Beijing Academy of Artificial Intelligence(北京智源人工智能研究院)
  • Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
  • YiXin-AILab(易信人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑