发表机构
University of Misan; Missan Oil Company(米桑大学; 米桑石油公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对伊拉克阿拉伯语评估缺失,提出国家基准Mizan,含MSA与伊拉克双赛道及六维度,经340条目构建,评估27个系统,发现MSA饱和而伊拉克赛道有区分度,并揭示过度拒绝问题。
AI 中文摘要
阿拉伯语大语言模型(LLM)评估已围绕现代标准阿拉伯语(MSA)趋于成熟:诸如Open Arabic LLM Leaderboard(OALL)、HELM Arabic和BALSAM等聚合排行榜在数十项MSA任务上对模型进行排名,前沿系统在这些任务上日益饱和。然而,伊拉克人实际使用的阿拉伯语方言,在这一基础设施中几乎不可见。我们推出Mizan(意为“天平”),这是伊拉克评估大语言模型在伊拉克阿拉伯语及伊拉克公民语境上表现的国家基准:包含一个MSA基线赛道和一个伊拉克赛道,后者涵盖六个维度(方言理解、方言生成、MSA-伊拉克语双向翻译、伊拉克特定知识、官方文档字段提取和安全性),由340个原创并经双重评审的条目构建,所有已发布分数均经过统计审计的答案位置和Wilson区间检验。对27个系统的试点评估——涵盖发布三天后的封闭前沿模型、各规模档次的开放权重模型,以及由商业、开放专用和主权系统组成的阿拉伯语三组合——得出四项发现。MSA赛道趋于饱和,而伊拉克赛道具有区分度,每个模型存在一致的14-18分差距,且领先者统计上并列。官方文档提取将所有系统限制在32-56分之间。阿拉伯语专用化表现为MSA专用化:两个专用阿拉伯语模型在伊拉克赛道上的得分低于同等规模的通才模型。最新模型系列中经过安全加固的层级,将无害的方言理解条目确定为政策违规而确定性拒绝,这是一种MSA基准无法显现的过度拒绝模式。该平台强制执行完整性协议,包括不可变快照、验证证书、人工发布门禁和公开撤回,这些都在本研究中得到实践。代码和公开开发集随论文一同发布。
英文摘要
Arabic large-language-model (LLM) evaluation has matured around Modern Standard Arabic (MSA): aggregated leaderboards such as the Open Arabic LLM Leaderboard (OALL), HELM Arabic, and BALSAM rank models across dozens of MSA tasks, and frontier systems increasingly saturate them. Dialectal Arabic, the language Iraqis actually speak, remains nearly invisible to this infrastructure. We introduce Mizan ("the balance"), Iraq's national benchmark for evaluating LLMs on Iraqi Arabic and the Iraqi civic context: an MSA baseline track paired with an Iraqi track across six axes (dialect comprehension, dialect generation, bidirectional MSA-Iraqi translation, Iraq-specific knowledge, official-document field extraction, and safety), built from 340 originally authored, dually reviewed items with statistically audited answer positions and Wilson intervals on every published score. A pilot evaluation of 27 systems, spanning closed frontier models three days after release, open weights across size tiers, and an Arabic trio of commercial, open-specialized, and sovereign systems, yields four findings. The MSA track saturates while the Iraqi track discriminates, with a consistent 14-18-point per-model gap and statistically tied leaders. Official-document extraction confines every system to 32-56. Arabic-focused specialization behaves as MSA specialization: two dedicated Arabic models score below a size-matched generalist on the Iraqi track. And the safety-hardened tier of the newest model family deterministically refuses innocuous dialect-comprehension items as policy violations, an over-refusal mode invisible to MSA benchmarks. The platform enforces an integrity protocol of immutable snapshots, verification certificates, a human publication gate, and public retraction, all exercised during this study. Code and the public development set accompany the paper.
Comments11 pages, 3 figures, 1 table. Live leaderboard: https://mizan-bench.onrender.com