arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11146cs.CL

低资源语言中跨语言安全的错觉

The Illusion of Cross-Lingual Safety in Low-Resource Languages

Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay, Maryam Ibrahim Mukhtar, Esmael Ahmed Abdu, Tassallah Abdullahi, Jessica Oparebea, Saminu Mohammad Aliyu, Idri… 展开作者

Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay, Maryam Ibrahim Mukhtar, Esmael Ahmed Abdu, Tassallah Abdullahi, Jessica Oparebea, Saminu Mohammad Aliyu, Idris Abdulmumin, Abubakar Juma Chilala, Nicholaus Dismas Ladislaus, Alfred Malengo Kondoro, Lemofouet Valdini Douglace, Shamsuddeen Hassan Muhammad, Seid Muhie Yimam

AI总结:

研究发现大型语言模型的跨语言安全迁移在四种非洲低资源语言中极为有限,现有多语言安全对齐仅停留在表面。

AI中文摘要:

大型语言模型(LLM)的安全对齐工作主要在英语中开展,假设这些安全措施可推广到多语言场景。但这一假设尚未得到充分探索,且在低资源语言中暴露出漏洞。我们使用新的安全数据集LoDNA(该数据集将字面翻译与文化本地化提示配对),研究了四种非洲语言(契维语、豪萨语、阿姆哈拉语、斯瓦希里语)的跨语言安全迁移。为超越基于生成的评估,我们提出一种潜在几何框架,用于探测LLM中隐藏状态的拒绝表征。实验结果显示,跨语言安全迁移极为有限;在大多数语言模型对中,有害提示仅保留不到10%的英语拒绝信号。字面提示与本地化提示在语义上对齐(余弦相似度为0.95-0.996),但跨层存在漂移,表明模型对相关概念进行了编码,但未将其路由至安全机制。这些发现表明,当前多语言安全对齐较为表面,为所研究的特定低资源语言中不存在通用、与语言无关的危害流形这一假设提供了有力反证。警告:本文包含可能具有冒犯性或有害性的示例数据。

英文摘要:

Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts. To move beyond generation-based evaluation, we propose a latent geometric framework that probes hidden-state refusal representations in LLMs. Our experimental results show that cross-lingual safety transfer is severely limited; harmful prompts retain less than 10% of the English refusal signal across most language-model pairs. Literal and localized prompts are semantically aligned (cosine 0.95-0.996) but drift across layers, suggesting models encode the concepts without routing them to safety mechanisms. These findings demonstrate that current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied. Warning: This paper contains example data that may be offensive or harmful.

↑