当安全“说”一种语言:大语言模型中安全-语言身份纠缠的机制分析
When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMs
浏览论文内容
中文总结 AI 辅助
本研究通过稀疏自编码器特征分析,揭示大语言模型中安全与语言身份的纠缠机制,发现安全相关特征具架构依赖性,为多语言安全干预提供机制解释。
中文摘要 AI 辅助
大语言模型(LLM)的安全对齐性能会随语言变化而下降,但导致这种不对称性的内部机制仍未被充分理解。因此,本研究采用稀疏自编码器(SAE)特征开展系统的多语言安全机制分析,涉及三个指令微调LLM、八种语言及所有模型层,研究了与有害、无害模型行为相关的残差流中的稀疏可解释方向。我们发现,安全相关特征的位置及跨层分布具有架构依赖性;此外,这些特征在几何上与语言身份纠缠,并呈现跨语言共享模式,即不同语言在模型深度和架构中以不同程度共享安全特征。这种安全-语言纠缠有直接影响:消融安全特征不仅会影响有害响应率,还会影响目标语言,干预程度可通过安全特征与语言特征的关系预测。本研究结果表明安全对齐的语言普遍性具有架构依赖性,并为多语言安全干预提供了机制解释。
英文摘要
Safety alignment of large language models (LLMs) degrades across languages, yet the internal mechanism driving this asymmetry remains poorly understood. Our work, therefore, presents a systematic mechanistic analysis of multilingual safety using sparse autoencoder (SAE) features, sparse interpretable directions in the residual stream associated with harmful and harmless model behavior across three instruction-tuned LLMs, eight languages, and all model layers. We observe that safety-relevant features are architecture-dependent in terms of where they are located and how they are distributed across layers. Additionally, they are geometrically entangled with language identity and exhibit cross-lingual sharing patterns, i.e., languages share safety features to varying degrees across model depths and architectures. This safety-language entanglement has direct consequences such that ablating safety features impacts not only harmful response rates but also target language, with the degree of intervention predicted by the relationship between safety and language features. Our findings qualify the language-universality of safety alignment as architecture-dependent and offer a mechanistic account of multilingual safety interventions.
发表机构
- L3S Research Center(L3S研究中心)
- Leibniz Universität Hannover(汉诺威莱布尼茨大学)
机构由 AI 辅助整理,请以论文原文为准。