发表机构
Elm Company(Elm公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对阿拉伯语事实知识评估不足的问题,提出动态评估框架AraDynFact,自动提取事实并生成问题,应用于阿拉伯语维基百科,验证了其与现有基准的高度相关性。
AI 中文摘要
随着大型语言模型(LLMs)在规模和能力上持续扩展,它们在阿拉伯语方面的熟练程度已取得显著进步。然而,一个关键缺口依然存在:它们对多样化阿拉伯语世界的事实知识掌握程度和文化敏感性在很大程度上仍未得到充分探索。当前的评估指标往往侧重于翻译或通用推理,未能捕捉阿拉伯文化中固有的丰富历史、社会和区域细微差别。此外,大多数基准测试依赖繁重的工作,在某些步骤中需要人工干预,使得知识覆盖范围的评估既昂贵又缓慢。为了解决这一不足,我们引入了AraDynFact,一个新颖的动态评估框架,旨在严格评估LLMs中嵌入的阿拉伯语事实知识。与静态基准不同,AraDynFact采用动态方法提取事实信息,并以快速、自动化的方式生成丰富且可回答的问题。我们将AraDynFact应用于阿拉伯语维基百科,并审计了多种最先进模型的性能,涵盖从以阿拉伯语为中心的专用LLMs到高资源通用LLMs。此外,我们发现与现有的手工制作的以阿拉伯语为中心的基准存在高度相关性,证实了我们动态方法的潜力。
英文摘要
As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.
CommentsAccepted to EMNLP 2026 Industry Track