arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22038cs.CL

QuranicMMLU:一个用于评估生成式AI解决方案在《古兰经》语言学知识上的认知感知基准

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

  • Qatar Computing Research Institute, Hamad Bin Khalifa University(哈马德·本·哈利法大学卡塔尔计算研究所)
  • Carnegie Mellon University Qatar(卡内基梅隆大学卡塔尔分校)
  • University of Doha for Science and Technology(多哈科技大学)
  • Qatar University(卡塔尔大学)
  • American University of Beirut(贝鲁特美国大学)
  • University of Birmingham(伯明翰大学)

机构由 AI 辅助整理,请以论文原文为准。

Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, … 展开作者

Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari

AI总结:

QuranicMMLU是一个基于五支柱语言学分类的基准,含980个人工审核问题,用于评估生成式AI在《古兰经》阿拉伯语上的表现,发现多项选择得分虚高,掩盖了开放式回答中的失败。

AI中文摘要:

我们推出了QuranicMMLU,这是一个用于在多个语言学复杂度维度上评估生成式AI在《古兰经》阿拉伯语上表现的基准。现有的《古兰经》基准侧重于通用问答和语义检索,并未探测特定的语言学能力,也未按认知需求和经文难度进行分层。我们构建了一个涵盖音系学、形态学、句法学、语义学和语用学五大支柱的《古兰经》分类体系,包含31个叶节点,覆盖从tajwīd(诵读规则)和词根-模式形态学到启示背景和章间连贯性等现象。对于每个叶节点,我们生成按Bloom认知水平和经文困惑度分层的问题,然后让LLM作为裁判独立回答并评分每个项目,并将注释路由至人工审查。最终数据集包含980个人工审查的问题,每个问题均以开放式和多项选择两种形式发布。我们在这些项目上对12个系统进行了基准测试,发现伊斯兰专用模型领先,但每个系统在多项选择准确率(平均84%)上的得分均高于开放式回答质量(平均60%):两种排名高度一致(Kendall's τ=0.73),但多项选择评分掩盖了仅在移除答案选项后才显现的失败。因此,QuranicMMLU为在《古兰经》领域评估阿拉伯语NLP提供了一个严谨的、基于语言学的框架。

英文摘要:

We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty. We construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf we generate questions stratified by Bloom's cognitive level and verse perplexity, then have LLM as a judge to independently answer and score every item and route the annotations to manual review. The resulting dataset comprises 980 human-reviewed questions, each issued in both open-ended and multiple-choice form. We benchmark 12 systems on these items and find that the Islamic-specialized model leads, yet every system scores higher on multiple-choice accuracy (average 84%) than open-ended answer quality (average 60%): the two rankings agree closely (Kendall's τ=0.73), but multiple-choice scoring hides failures that surface only once answer choices are removed. QuranicMMLU thus offers a rigorous, linguistically grounded framework for evaluating Arabic NLP in the Quranic domain.

补充信息

↑