arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArabicDialectSafety:一种面向阿拉伯语内容安全分类的方言感知基准

ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

Wajdi Zaghouani, Md. Rafiul Biswas, Kholoud Khalil Aldous, Mabrouka Bessghaier

arXiv 2608.01291首次发表:更新:

发表机构

Northwestern University in Qatar; Hamad Bin Khalifa University(卡塔尔西北大学; 哈马德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建了覆盖6种阿拉伯语变体的安全基准数据集,提出双任务评估框架,发现微调后的MARBERTv2性能最优,低资源马格里布方言仍存差距,前沿大语言模型不安全生成率低于5%。

AI 中文摘要

我们推出ArabicDialectSafety,这是一个人工整理的阿拉伯语安全数据集,包含25071条提示,覆盖六种阿拉伯语变体:现代标准阿拉伯语、叙利亚语、埃及语、阿尔及利亚语、巴勒斯坦语和摩洛哥语。该数据集标注了方言标签和七个细粒度危害类别。我们引入了一个双任务评估框架,用于跨方言的安全/不安全二元检测和细粒度危害分类。对七种监督模型和生成模型进行基准测试后,我们发现微调后的MARBERTv2表现最强,二元分类的Macro-F1分数为0.95,细粒度分类的Macro-F1分数为0.90,显著优于提示式前沿大语言模型(包括阿拉伯语专用模型)。我们的分析表明,方言条件在表示层面整合时效果最佳,而低资源马格里布方言仍存在显著性能差距。我们进一步评估了七种前沿大语言模型作为响应生成器,针对有害阿拉伯方言提示,所有模型的不安全生成率均低于5%。我们将在论文接收后发布数据集和代码,以支持未来方言感知的阿拉伯语安全评估研究。警告:本文包含有害及潜在冒犯性内容示例,仅用于研究目的。

英文摘要

We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with dialect labels and seven fine-grained harm categories. We introduce a dual-task evaluation framework for binary safe/unsafe detection and granular harm classification across dialects. Benchmarking seven supervised and generative models, we find that fine-tuned MARBERTv2 achieves the strongest performance, with Macro-F1 scores of 0.95 for binary classification and 0.90 for granular classification, substantially outperforming prompted frontier LLMs, including Arabic-specialized models. Our analyses show that dialect conditioning is most effective when integrated at the representation level, while significant performance gaps remain for low-resource Maghrebi dialects. We further evaluate seven frontier LLMs as response generators on harmful dialectal Arabic prompts and observe unsafe generation rates below 5 percent across models. We release the dataset and code upon acceptance to support future research on dialect-aware Arabic safety evaluation. Warning: This paper contains examples of harmful and potentially offensive content included solely for research purposes.

Comments13 pages, 2 figures, 9 tables

Journal refEMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑