SurakshaEval:面向多语言大语言模型的印度语安全基准
SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
浏览论文内容
中文总结 AI 辅助
针对现有LLM安全评估数据集忽视印度语言文化安全风险的问题,本文推出SurakshaEval基准,测试发现多语言LLM在印度语场景下安全表现不足,凸显了适配区域数据的安全评估框架的必要性。
中文摘要 AI 辅助
现有的大语言模型(LLM)安全评估数据集主要聚焦于英语和西方语境,往往忽视了其他语言中存在的语言多样性和基于文化的安全风险。为解决这一缺口,我们推出SurakshaEval,这是一个新颖的安全基准,由涵盖现实场景的人工编写提示词组成,明确面向十种主要印度语言——阿萨姆语、孟加拉语、古吉拉特语、印地语、卡纳达语、马拉雅拉姆语、马拉地语、旁遮普语、泰米尔语和泰卢固语,同时也包含英语。SurakshaEval既包含印度通用的通用提示词,也包含捕捉本地化社会文化敏感性的特定地区和语言的提示词。我们在SurakshaEval上对广泛的最先进LLM进行基准测试,建立了安全性能基线,并识别出反复出现的失败模式,包括过度弃权(不执行)、未检测到隐含偏见,以及在区域敏感场景中情境意识不足。我们的结果表明,即使是强大的多语言LLM在使用印度语(尤其是原生脚本)运行时,也难以可靠满足细微的安全要求。这些发现凸显了迫切需要纳入特定地区数据和结构化评估协议的安全评估框架,以开发和部署安全、符合伦理且与多元社会价值观一致的AI系统。我们的代码和数据可在此https URL获取。警告:本文包含可能具有冒犯性或不安全的文本。
英文摘要
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, along with English. SurakshaEval includes both generic prompts common across India and region- and language-specific prompts that capture localized sociocultural sensitivities. We benchmark a broad range of state-of-the-art LLMs on SurakshaEval, establish baseline safety performance, and identify recurring failure modes, including over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings. Our results show that even strong multilingual LLMs struggle to reliably meet nuanced safety requirements when operating in Indic languages, particularly in native scripts. These findings highlight the urgent need for safety evaluation frameworks that incorporate region-specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. Our code and data are available at https://github.com/debobanerjee/SurakshaEval. Warning: This paper contains text that may be offensive or unsafe.
发表机构
- Inception42
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
- University of Michigan(密歇根大学)
- Institute of Computer Science Artificial Intelligence & Technology(计算机科学人工智能与技术研究院)
- International Institute of Information Technology Hyderabad(海得拉巴国际信息技术学院)
- Ministry of Electronics and IT India(印度电子与信息技术部)
- Indian Institute of Technology Kharagpur(印度理工学院卡拉格普尔分校)
机构由 AI 辅助整理,请以论文原文为准。