低资源语言下的大语言模型安全对齐:系统文献综述
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review
浏览论文内容
中文总结 AI 辅助
本文通过PRISMA 2020方法学开展系统文献综述,分析50项相关研究,揭示低资源语言下LLM存在多语言安全差距,提出安全对齐分类,并给出未来研究方向。
中文摘要 AI 辅助
大语言模型(LLMs)在安全对齐方面已取得显著进展,但其在低资源语言和多语言场景中的安全保障仍远弱于高资源语言。本文采用PRISMA 2020方法学,对低资源语言下的LLM安全对齐开展系统文献综述(SLR)。从Semantic Scholar、arXiv和OpenAlex中识别出约1500篇论文后,筛选并分析了50项相关研究。本综述围绕四大主题展开:安全对齐方法、多语言安全风险、评估基准及跨语言可迁移性。我们进一步基于三种适配机制提出安全对齐方法的分类:数据适配、目标优化及机制对齐。现有文献表明,翻译后的英语基准无法充分代表植根于文化的危害,多语言模型更易受到跨语言越狱攻击、语码切换攻击,以及在代表性不足的语言中出现安全性能下降。这些失效由多个关键因素导致,包括多语言预训练覆盖不均、原生语言偏好数据不足、安全表示的迁移效果差,以及缺乏具有文化意识的评估框架。综述还指出,许多低资源语言,尤其是非洲语言,可用的安全基准少于其他多语言区域。总体而言,研究结果揭示了持续存在的多语言安全差距,表明未来进展需要以文化为基础的基准、参与式数据收集、均衡的多语言预训练,以及可扩展的多语言对齐方法。
英文摘要
Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper, we conduct a Systematic Literature Review (SLR) of LLM safety alignment in low-resource languages by adopting the PRISMA 2020 methodology. Out of roughly 1,500 papers identified from Semantic Scholar, arXiv, and OpenAlex, 50 relevant studies have been selected and analyzed. Our review is organized around four themes: safety alignment methods, multilingual safety risks, evaluation benchmarks, and cross-lingual transferability. We further propose a taxonomy of safety alignment approaches based on three adaptation mechanisms: data adaptation, objective optimization, and mechanistic alignment. Across literature, translated English benchmarks fail to sufficiently represent culturally rooted harms, and multilingual models are more vulnerable to cross-lingual jailbreaks, code-switching attacks, and safety degradation in underrepresented languages. These failures are driven by several key factors, including uneven multilingual pre-training coverage, insufficient native-language preference data, poor transfer of safety representations, and a lack of culturally aware evaluation frameworks. The review also notes that many low-resource languages, especially African languages, have fewer safety benchmarks available than other multilingual regions. Overall, the results reveal a persistent multilingual safety gap, and suggest that future progress will require culturally grounded benchmarks, participatory data collection, balanced multilingual pre-training, and scalable multilingual alignment methods.
发表机构
- African Institute for Mathematical Sciences (AIMS)(非洲数学科学研究所)
- Bayero University Kano(卡诺贝罗大学)
- Brown University(布朗大学)
- Centrum Wiskunde & Informatica(数学与计算机科学中心)
- University of Pretoria(比勒陀利亚大学)
- University of Hamburg(汉堡大学)
- Imperial College London(伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。