发表机构
University of Göttingen; German State Police NRW(哥廷根大学; 北莱茵-威斯特伐利亚州德国警察)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出BabelSteering激活引导方法,利用英语安全监督的拒绝方向实现多语言安全对齐,可提升多语言有害请求拒绝率且任务效用损失极小,还引入多语言翻译评估流程。
AI 中文摘要
大型语言模型(LLMs)在全球高风险场景中部署,但大多数安全研究与对齐工作仍集中在英语领域。因此,使用其他语言与LLMs交互的用户,即便在执行类似敏感任务时依赖相同系统,也可能遭遇较弱的安全防护。本研究探讨从高资源语言(如英语)中学习到的安全信号是否能提升多语言安全性。我们提出BabelSteering,这是一种激活引导方法,作为轻量级推理时干预手段,利用从英语安全监督中推导的拒绝方向实现跨语言泛化。我们的评估涵盖8种语言,联合衡量有害请求的拒绝、过度拒绝及通用任务效用。结果显示,BabelSteering可提升各语言对有害请求的拒绝率,仅带来边际或无任务效用下降,但对伪有害提示的拒绝有所增加。例如,对于Gemma 7B,我们观察到各语言对有害提示的拒绝率平均提升11个百分点(pp),其中孟加拉语等个别语言提升17个百分点,Global MMLU上无效用损失,伪有害提示拒绝率平均提升13个百分点。我们还引入多语言翻译-评估流程,以推动跨语言安全干预的后续研究。总体而言,我们的发现表明,激活引导或可成为将英语衍生安全信号扩展至其他语言的实用低成本机制。警告:本文包含含不安全内容的示例。
英文摘要
Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users interacting with LLMs in other languages may encounter weaker safeguards despite relying on the same systems for similarly sensitive tasks. In this work, we investigate whether safety signals learned from a high-resource language, like English, can improve multilingual safety. We propose BabelSteering, an activation steering method that acts as a lightweight inference- time intervention, using refusal directions derived from English safety supervision to generalize across languages. Our evaluation includes eight languages and jointly measures refusal of harmful requests, over-refusal, and general task utility. The results show that BabelSteering increases the refusal of harmful requests across languages, with only a marginal to no reduction in task utility but with some increase in refusal of pseudo-harmful prompts. For example, for Gemma 7B, we see an average increase in the refusal of harmful prompts across languages of 11 percentage points (pp), with individual languages like Bengali seeing an increase of 17 pp, with no loss of utility on Global MMLU, while pseudo-harmful refusals increase by 13 pp on average. We also introduce a multilingual translation-and-evaluation pipeline to facilitate future work on cross-lingual safety interventions. Overall, our findings suggest that activation steering may provide a practical, low- cost mechanism for extending English-derived safety signals to other languages. Warning: this paper contains examples with unsafe content