SafeEvo:解析语言模型的安全对齐机制与演化
SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
浏览论文内容
中文总结 AI 辅助
SafeEvo从电路视角揭示预训练模型中的弱拒绝电路及其演化,提出安全电路对齐(SCA)方法,在三个LLM上实现更强对齐、更少过度拒绝和更好效用。
中文摘要 AI 辅助
安全可解释性将大语言模型(LLM)对齐的研究从由数据或算法驱动的行为约束推进到对内部机制的更深入理解。然而,现有工作主要聚焦于对齐后的安全相关表示、注意力头或神经元,而在很大程度上忽略了预训练模型中的安全机制及其在对齐检查点间的演化。为解决这一问题,我们提出了SafeEvo,一个从电路(LLM的稀疏子图)视角出发的可解释性框架。SafeEvo首先应用基于优化的提取算法来识别预训练基础LLM中能够独立表达拒绝行为的弱拒绝电路。因果消融这些电路会完全消除基础模型对有害输入的拒绝。随后,SafeEvo追踪拒绝电路在连续对齐检查点间的演化,发现其结构逐步变化,表明对齐税可能源于拒绝电路更新对效用相关参数的影响。为验证这一点,SafeEvo引入了安全电路对齐(SCA),将安全更新限制在拒绝电路内。在三个LLM和两种对齐算法上的实验表明,平均而言,SCA在三个方面优于普通对齐:(1)更强的对齐性,将有害性评分降低63.21%;(2)更少的过度拒绝,良性查询的拒绝率降低58.44%;(3)更好的效用,保留原始模型能力的99.58%。
英文摘要
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model's refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf{(1) stronger alignment}, lowering harmfulness score by 63.21\%; \textbf{(2) less over-refusal}, yielding a 58.44\% decrease in refusal rates for benign queries; and \textbf{(3) better utility}, retaining 99.58\% of the original model capabilities.
发表机构
- The University of Hong Kong (HKU)(香港大学)
- Chinese Academy of Sciences (CAS)(中国科学院)
- Information Engineering University (IEU)(信息工程大学)
- Nanyang Technological University (NTU)(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。