arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03781cs.CLcs.AI

IndicSafeEval:多语言劝诱式越狱攻击下大语言模型的安全鲁棒性

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

  • Indian Institute of Technology Jodhpur(印度焦特布尔印度理工学院)
  • King’s College London(伦敦国王学院)
  • Indian Institute of Technology Patna(印度巴特那印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal

AI总结:

本文提出针对印度语言的多语言劝诱式越狱评估框架IndicSafeEval,通过7200条对抗性提示评估开源LLMs,发现其安全表现受语言、劝诱策略及风险类别影响,凸显多语言安全评估框架的必要性。

AI中文摘要:

大语言模型(LLMs)在多语言场景中的应用日益广泛,但其安全性评估仍主要以英语为核心,这限制了我们对齐失败在低资源语言及文化多样语言中的表现形式的理解。本文提出了IndicSafeEval,这是一个针对印度语言的基于劝诱的越狱评估框架。该基准将10类安全关键内容类别与6类类人劝诱策略相结合,覆盖印地语、孟加拉语、马拉地语、旁遮普语共4种不同的印度语言,生成了7200条对抗性提示。我们对多个开源LLMs开展了系统性黑盒评估,以考察其安全行为在不同语言、劝诱策略及风险类别下的差异。分析结果显示,模型在不同语言和提示风格下的安全表现并不一致,安全性能强烈依赖于使用的语言以及通过劝诱线索表述请求的方式;同时,不同风险类别的脆弱性水平存在差异,部分类型的有害内容更易受基于劝诱的越狱攻击影响。这些发现揭示了当前以英语为中心的安全评估存在的重要局限性,强调了构建多语言、感知劝诱的基准框架以更准确评估现实世界LLM安全性的必要性。本文的实现代码可通过指定URL获取。警告:本文包含可能具有冒犯性或有害性的示例数据。

英文摘要:

Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such as Hindi, Bengali, Marathi and Punjabi, resulting in 7,200 adversarial prompts. We conduct a systematic black-box evaluation of several open-source LLMs to examine how their safety behaviour varies across languages, persuasion strategies, and risk categories. Our analysis shows that the model does not behave equally safely across all languages and prompt styles. Instead, safety performance depends strongly on both the languages used and the way a request is phrased using persuasive cues. We further observe that different risk categories exhibit different levels of vulnerability, with some types of harmful content being significantly more susceptible to persuasion-based jailbreaks than others. These findings reveal important limitations of current safety evaluations, which are largely English-centric, and underscore the need for multilingual and persuasion-aware benchmarking frameworks to more accurately assess real-world LLM safety. Our implementation is available at https://github.com/MonSaikat/IndicSafeEval. Warning: this paper contains example data that may be offensive or harmful.

补充信息

↑