arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

干预干扰反映模型的默认倾向,而非行为方向

Steering Interference Reflects the Model's Defaults, Not the Behavior Directions

Srikanth Malla, Chiho Choi, Joon Hee Choi

arXiv 2609.06951首次发表:更新:

发表机构

Samsung Semiconductor US(三星半导体美国公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现激活干预的副作用由模型默认倾向决定而非行为方向,解耦行为方向无法实现模块化控制。

AI 中文摘要

激活干预有望实现对语言模型行为的模块化控制:诸如礼貌等行为对应于模型激活中的一个方向,在生成过程中添加该方向应能开启该行为而不影响其他方面。然而事实并非如此。我们探究是什么决定了哪些其他行为会随之改变以及改变的程度,发现起决定作用的是模型本身,而非被干预的行为。一次干预会使模型趋向于它本就偏好的一小组行为,主要是拒绝、谄媚和诗意化,且无论干预何种行为,这一组行为都大致相同。在24种行为和十个指令微调模型上的三项结果支持了这一结论,所有效果均由语言模型评判器从生成文本中读出,而非通过探针获取。这种读出方式至关重要:全部24种行为均可线性解码,但只有20种会改变模型的实际输出。第一,一个不携带任何行为内容的方向,仅在所加向量的大小上与真实干预匹配,其引发的行为变化及顺序与真实干预相同,却不会产生任何需要特定方向才能实现的行为。第二,多数干扰是单向的,因此不可能是两个方向之间的重叠:干预粗俗行为会使模型变得有毒,而干预毒性却不会影响粗俗行为。第三,在完全排除某一行为的情况下,基于其他行为测量的几何特性几乎无法解释该行为所参与的干扰。这一结论在全部十个模型上均成立,向默认倾向的拉拽在参数低于100亿时最强,并在每个模型系列的最大版本中减弱。将干预视为一种扰动、其终点由模型自行确定,意味着解耦行为方向本身并不能使干预实现模块化。

英文摘要

Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a probe. That readout matters: all 24 behaviors are linearly decodable, but only 20 change what the model writes. First, a direction carrying no behavioral content, matched to a real steer only in the size of the vector it adds, moves the same behaviors in the same order as real steers do, while producing none of the behaviors that need a specific direction. Second, most interference runs one way, so it cannot be an overlap between two directions: steering profanity makes the model toxic, while steering toxicity leaves profanity untouched. Third, with a behavior held out entirely, geometry measured on the others explains almost none of the interference it takes part in. The account holds on all ten models, the pull toward defaults strongest below 10B parameters and weakening in each family's largest. Reading a steer as a perturbation whose endpoint the model fixes implies that disentangling behavior directions cannot by itself make steering modular.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑