定位并引导注意力之外的拒绝行为
Locating and Steering Refusal Beyond Attention
浏览论文内容
中文总结 AI 辅助
该研究发现不同架构的语言模型共享拒绝行为的表示方向,通过在新架构写入位置重新估计该方向,可将安全工具迁移至多种架构,降低越狱成功率。
中文摘要 AI 辅助
语言模型的拒绝行为发生在模型内部的哪个位置?当模型架构发生变化时,该位置是否会改变?在Transformer架构中,拒绝行为由残差流中的单个方向控制,这一发现是当前安全与可解释性工具所依赖的基础。状态空间模型(SSMs)通过循环更新而非注意力机制来传递信息,与Transformer没有相同的token混合机制。那么相同的安全表示能否在这种架构转变中保留,还是必须针对每种架构重新发现?答案是可以保留。仅通过一次刚性旋转(该旋转只能重新定向空间而不能重塑空间),即可将一个模型的表示空间与另一个模型的表示空间对齐,因此两个模型确实共享该表示。在Transformer上训练的有害探测模型随后可标记SSM的有害输入;移除该对齐方向会使模型对原本会拒绝的攻击做出回应,而相同大小的随机方向则效果弱得多。架构特定的并非该方向的引导位置,而是其读取位置。每层计算出一个新的输出,随后被添加到残差流中;在添加操作之前,在该输出(即写入位置)处可清晰读取有害信息。一项控制干预强度固定的实验表明,关键因素是方向的估计位置,而非其应用位置。通过检测器触发的门控应用该方向,可降低我们测试的所有四种架构家族(SSM、Transformer、循环、混合)的越狱成功率,且在SSM上,即使攻击者针对防御调整提示,该方向仍能发挥作用。该门控仅遵循简单规则:每当相同检测器触发时,返回固定的弃权(不执行);因此跨架构迁移的是方向本身,而非防御强度。因此,基于拒绝行为构建的安全工具,只需在新架构的写入位置重新估计该方向,即可迁移到该新架构,无需重建工具。
英文摘要
Where inside a language model does refusal live, and does that place change when the architecture does? In a transformer, refusal is governed by a single direction in the residual stream, a finding that safety and interpretability tooling now depend on. State-space models (SSMs) route information through a recurrent update instead of attention, sharing no token-mixing mechanism with a transformer. Does the same safety representation survive this shift, or must it be rediscovered per architecture? It survives. A single rigid rotation, which can only reorient a space and not reshape it, aligns one model's representation space with another's, so the two genuinely share the representation. A harm probe trained on a transformer then flags an SSM's harmful inputs, and removing the aligned direction makes a model answer attacks it would otherwise refuse, while a random direction of the same size does far less. What is architecture-specific is not where the direction is steered but where it must be read. Each layer computes a fresh output that is then added into the residual stream, and harm is cleanly readable at this output, the write site, before the addition. A control that holds the intervention's strength fixed shows that what matters is where the direction is estimated, not where it is applied. Applied through a detector-triggered gate, this direction lowers jailbreak success in all four architecture families we test (SSM, transformer, recurrent, hybrid), and on the SSM it holds against an attacker that tunes its prompt against the defense. The gate only matches a trivial rule that returns a fixed refusal whenever the same detector fires, so what transfers across architectures is the direction itself, not defense strength. Safety tooling built on refusal therefore ports to a new architecture by re-estimating the direction at that architecture's write site, not by rebuilding it.
发表机构
- Department of Computer Science and Engineering, Indian Institute of Technology Madras(印度马德拉斯理工学院计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。