arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

安全对齐错觉:大语言模型中的跨语言安全差距

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

Namya Bhatnagar

arXiv 2608.18131首次发表:更新:

AI 中文总结

针对 LLMs 安全对齐以英语为中心导致的跨语言偏见问题,提出多语言评估基准 INCLUDE,评估十款 LLMs 后发现孟加拉语在开源模型中偏见最高、英语在开源与闭源模型中偏见表现反转的现象。

AI 中文摘要

当前大语言模型(LLMs)的安全对齐训练严重以英语为中心。当这类安全过滤器在非英语语言中失效时,后果会直接面向用户:语音助手和口语对话系统可能产生强化刻板印象的输出,绕过以英语为核心的标准安全对齐,并将有害偏见传播到非英语社区。对于部署在语言多样的印度人口中的口语技术而言,这是一种关键的失效模式。为解决这一跨语言差距,我们提出 INCLUDE(印度文化偏见理解与检测基准,Indian Cultural Lens for Understanding and Detecting Embedded Biases),这是一个旨在量化印度中心社会文化偏见的多语言评估基准。INCLUDE 包含 2604 个提示,涵盖六种提示语言:英语、印地语、孟加拉语、马拉地语、泰米尔语和印式英语(印地语-英语混合语)。我们针对该基准评估了十个开源和闭源 LLMs,分析了 14988 个偏见分数。统计结果揭示两个关键发现:第一,孟加拉语在开源模型中产生最高的平均偏见分数;第二,英语表现出显著反转,在开源模型中产生最低偏见,但在闭源模型中产生最高偏见。

英文摘要

Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken dialogue systems may produce stereotype-reinforcing outputs, bypassing the standard English-focused safety alignments and propagating harmful bias to non-English speaking communities. For spoken language technologies deployed across India's linguistically diverse population, this represents a critical failure mode. To address this cross-lingual gap, we introduce INCLUDE (Indian Cultural Lens for Understanding and Detecting Embedded Biases), a multilingual evaluation benchmark designed to quantify Indian-centric socio-cultural biases. INCLUDE consists of 2,604 prompts spanning six prompt languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish (Hindi-English code-mix). We evaluate ten open- and closed-source LLMs against this benchmark, analyzing 14,988 bias scores. Our statistical results reveal two key findings. First, Bengali yielded the highest average bias score in open-source models. Second, English demonstrated a notable reversal, producing the lowest bias in open-source models but the highest bias in closed-source models.

Comments7 pages, 8 figures, submitted to IEEE SLT (Spoken Language Technology) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑