AI 中文总结
本研究推出首个人工策划的阿拉伯语红队测试基准ASAS,评估7款领先阿拉伯LLM的安全性,发现多数模型对50%不安全提示防御失效,自动安全评判器表现逊于人工,为阿拉伯LLM安全提供文化适配的测试方案。
AI 中文摘要
随着大型语言模型(LLMs)在阿拉伯语地区的应用不断增长,确保其安全性和文化适配性愈发关键。然而,阿拉伯LLM的安全性研究仍未得到充分探索,尤其是在对抗性评估场景中。我们推出阿拉伯安全指数(ASAS),这是首个完全由人工策划的用于LLM红队测试的阿拉伯语基准。ASAS包含801条提示,覆盖8个安全类别和8种攻击策略,理想响应采用现代标准阿拉伯语(MSA)。我们对7款具备阿拉伯语能力的领先模型开展红队评估,包括GPT-4o、Claude 3.7 Sonnet,以及ALLAM、FANAR等区域模型。人工标注者采用结构化4点安全量表对响应进行评分,结果显示多数模型对50%的不安全提示无法防御。研究发现,在武器、非法物质等高危害类别中存在重大安全缺口,直接攻击和基于混淆的攻击效果最佳。结果还表明,语言适配性无法在不同语言间轻易迁移,且自动安全评判器(如GPT-4o)的表现逊于人工标注者。ASAS提供了符合文化背景的基准和红队测试方案,以推动阿拉伯LLM安全性的进步。
英文摘要
As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.
Comments13 pages, 5 figues, 3 tables