《翻译中的隐秘危害:利用“乌尔语未检出”分数测量大型语言模型仇恨言论检测中的跨文字安全不一致性》
'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection
浏览论文内容
中文总结 AI 辅助
该研究针对乌尔语在LLM安全评估中的缺失问题,测试五款LLM在多类乌尔语相关数据集上的表现,发现其跨文字安全不一致性,且开放权重模型问题更显著。
中文摘要 AI 辅助
乌尔语是全球使用人数第十多的语言,使用者达2.46亿,但在主流大型语言模型(LLM)安全评估以及历时九年的WOAH(仇恨言论工作坊)相关研究中几乎完全缺席。为探究这种缺失是否会对内容审核可靠性产生可测量的影响,研究人员对GPT-4o、Claude Sonnet 4.5、Gemini 2.5 Flash、Qwen-2.5和Llama-3.1这五款大型语言模型,在涵盖 Nastaliq 乌尔语、罗马乌尔语、英语及乌尔语-英语代码切换文本的六个数据集上开展了测试。在五款乌尔语文字数据集上,原文字与英文翻译分类之间的标签不稳定性范围为15.9%(Gemini 2.5 Flash)至31.6%(Qwen-2.5);“乌尔语未检出”率(即内容在英文翻译中被标记为有害,但在原文字中被判定为正常)范围为2.4%至9.9%(中位数为4.3%)。通过ACL Anthology API对九届ALW/WOAH会议的全部205篇论文进行完整枚举后确认,整个研究期间没有一篇专门针对乌尔语的论文。结果表明,当前LLM在乌尔语的不同文字变体上提供的安全保障存在不均等性,规模较小的开放权重模型的不稳定性和未检出危害率明显高于前沿闭源模型。
英文摘要
Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings. To investigate whether this absence has measurable consequences for content moderation reliability, five large language models, GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5, and Llama-3.1, were tested across six datasets spanning Nastaliq Urdu, Roman Urdu, English, and code-switched Urdu-English. Across the five Urdu-script datasets, label instability between original-script and English-translation classification ranged from 15.9% (Gemini 2.5 Flash) to 31.6% (Qwen-2.5), with a 'Missed-in-Urdu' rate, content flagged as harmful in English translation but passed as normal in the original script, ranging from 2.4% to 9.9% (median 4.3%). A complete enumeration of all 205 papers across nine ALW/WOAH editions via the ACL Anthology API confirms zero dedicated Urdu papers across the entire period. Results indicate that current LLMs provide uneven safety assurance across Urdu's script varieties, with smaller open-weight models showing substantially higher instability and missed-harm rates than frontier closed models.