发表机构
Mohamed bin Zayed University of Artificial Intelligence; New York University Abu Dhabi; IBM Research AI; Qassim University(穆罕默德·本·扎耶德人工智能大学; 纽约大学阿布扎比分校; IBM研究院人工智能部门; 卡西姆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对阿拉伯语方言NLP资源不足的问题,推出首个涵盖5种方言的大规模MRC基准EDRAC,测试发现现有评估指标存在局限,为相关研究提供支撑。
AI 中文摘要
与现代标准阿拉伯语(MSA)相比,方言阿拉伯语(DA)的资源仍然不足,尤其是在机器阅读理解(MRC)和问答(QA)领域。现有的阿拉伯语QA基准主要聚焦于正式书面的MSA或多项选择QA,对自然口语方言的覆盖有限。为此,本文旨在弥合这一差距,推出EDRAC——首个大规模阿拉伯语方言机器阅读理解(MRC)和生成式问答基准,涵盖埃及、摩洛哥、阿联酋、叙利亚和沙特阿拉伯五种主要方言。EDRAC包含499篇源自自然口语互动的段落,以及通过人类与大语言模型(LLM)协作流程生成的4977组对应QA对,该流程结合了迭代生成、LLM作为评判者的评估以及人工验证。研究人员使用EDRAC对阿拉伯语中心型和多语言大语言模型(LLM)进行基准测试,采用词汇和语义指标。结果显示,语义答案质量与方言保真度之间存在显著差距,凸显了现有方言阿拉伯语生成评估指标的局限性。EDRAC为未来阿拉伯语方言自然语言处理(NLP)研究提供了一个现实且具有挑战性的MRC基准。
英文摘要
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spoken dialects. Here, we aim to bridge this gap. We introduce EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic. EDRAC contains 499 passages derived from naturally occurring spoken interactions and 4,977 corresponding QA pairs generated through a human--LLM collaborative pipeline combining iterative generation, LLM-as-a-judge evaluation, and human verification. We benchmark Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics. Our results reveal substantial gaps between semantic answer quality and dialectal fidelity, highlighting the limitations of existing evaluation metrics for dialectal Arabic generation. EDRAC provides a realistic and challenging MRC benchmark for future research on dialectal Arabic NLP.