发表机构
University of Illinois Urbana-Champaign; IBM Research(伊利诺伊大学厄巴纳-香槟分校; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究语言模型拒绝行为的引导维度,发现安全对齐触发的拒绝可1维引导,而一般情境拒绝需更高维度,揭示其激活几何结构复杂且依赖情境。
AI 中文摘要
现有的激活引导方法通常假设一个高层概念可以通过单一的引导方向来中介。为了支持这一点,应实现两个互补的干预措施:加法引导应诱导目标行为,而方向性消融应抑制它。然而,行为可能占据比单一方向更丰富的激活几何结构,且语义相似的行为可能由不同的方向表示。在这项工作中,我们研究了有多少方向能够可靠地控制两种不同类型的拒绝行为:由安全对齐触发的拒绝和一般情境下的拒绝。鉴于单一引导向量的表达能力有限,我们研究了引导子空间的一般设置,并引入了引导维度的概念,即可靠控制行为所需的最小子空间维度。我们通过充分性维度之前的(单调)改进以及超过该维度后的饱和,来刻画覆盖目标行为全部范围的充分引导子空间。实验上,我们发现由安全对齐触发的拒绝是1维可引导的,而多个不同的引导方向可以实现相当的控制效果。相比之下,一般情境下的拒绝表现出显著更丰富的激活几何结构,即使5维引导子空间也无法可靠地捕获其全部可引导变化。我们的结果表明,拒绝行为背后的激活几何结构高度依赖于情境,并且可能比单一线性引导方向复杂得多。
英文摘要
Existing activation steering methods often assume that a high-level concept can be mediated by a single steering direction. To support this, two complementary interventions should be achieved: additive steering should induce the target behavior, while directional ablation should suppress it. Yet behaviors may occupy richer activation geometries beyond a single direction, and semantically similar behaviors may be represented by distinct directions. In this work, we study how many directions can reliably control two different types of refusal behaviors: refusal triggered by the safety alignment and refusal in general contexts. Given the limited expressive capability of a single steering vector, we study the general setting of steering subspaces and introduce the notion of steering dimensionality as the minimum subspace dimensionality required to reliably control a behavior. We characterize sufficient steering subspaces that cover the full extent of the target behavior through both the (monotonic) improvement before the sufficient dimensionality, and the saturation beyond it. Empirically, we find that refusal triggered by safety alignment is 1-dim steerable, while multiple distinct steering directions can achieve comparable control. In contrast, refusal in general contexts exhibits substantially richer activation geometry where even 5-dim steering subspaces fail to reliably capture its full steerable variation. Our results reveal that the activation geometry underlying refusal is highly context-dependent and can be substantially more complex than a single linear steering direction.
CommentsInterpScience Workshop at NeurIPS 2026