大语言模型拒绝回答中的对称性破缺:答案释放比拒绝恢复更具局部性
Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration
浏览论文内容
中文总结 AI 辅助
本研究通过受控保留设置揭示大语言模型拒绝回答存在对称性破缺:答案释放具高度局部性,拒绝恢复需更广干预,拒绝并非简单对称开关,对安全审计具启示。
中文摘要 AI 辅助
当语言模型拒绝回答提示时,目前尚不清楚正确答案是从其内部表征中被擦除,还是仅在输出层被抑制。我们采用受控保留设置来研究这一机制,该设置能为双向激活修补生成完全匹配的回答与拒绝轨迹。我们在匹配的因果干预下发现了干预局部性的因果不对称性,将其命名为对称性破缺。即使模型生成了干净的拒绝,仍可从其隐藏状态中线性恢复出正确答案。此外,释放该保留的答案是一个高度局部化的操作,仅需单个位置的修补即可完成。相反,反向操作并不具有同等局部性:重新施加抑制需要跨多个位置的更广泛干预,而组装连贯的拒绝序列则更为困难。我们进一步证明,虽然平均答案到拒绝的位移向量标记了这些状态之间的几何差异,但它无法作为行为之间可靠、可逆的线性控制开关。综合来看,我们的发现表明,拒绝并非简单的对称开关。对于安全和审计而言,这意味着探测可恢复性可能会高估实际的行为控制,且定位与拒绝相关的方向并不可靠地赋予从回答转向连贯拒绝的能力。
英文摘要
When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withhold setting, which yields perfectly matched answering and refusal trajectories for bidirectional activation patching. We uncover a causal asymmetry in intervention locality under matched causal interventions, which we term broken symmetry. Even when a model generates a clean refusal, the correct answer remains linearly recoverable from its hidden states. Furthermore, releasing this withheld answer is a highly local operation, requiring only a single-position patch. Conversely, the reverse operation is not equally local: reimposing suppression requires broader interventions across multiple positions, and assembling a coherent refusal sequence is more difficult still. We further demonstrate that while an average answer-to-refusal displacement vector marks the geometric difference between these states, it fails to act as a reliable, reversible linear control toggle between behaviours. Taken together, our findings show that refusal does not function as a simple symmetric switch. For safety and auditing, this implies that probe recoverability can overestimate true behavioural control, and locating refusal-relevant directions does not reliably grant the ability to steer a model from answering to coherent refusal.
发表机构
- University of Manchester(曼彻斯特大学)
- Shanghai University of Finance and Economics(上海财经大学)
机构由 AI 辅助整理,请以论文原文为准。