ARAFA:一个由大语言模型生成的阿拉伯语事实核查数据集
ARAFA: An LLM-Generated Arabic Fact-Checking Dataset
浏览论文内容
中文总结 AI 辅助
本文提出Arafa,首个大规模阿拉伯语事实核查数据集,通过LLM自动化三步流水线构建,含181,976个声明-证据对,并验证了其质量与微调模型的性能。
中文摘要 AI 辅助
自动事实核查在阿拉伯语自然语言处理中因其数据集和资源的稀缺而构成重大挑战。在本文中,我们介绍了Arafa,一个用于现代标准阿拉伯语事实核查的新的大型数据集,该数据集通过一个利用大语言模型(LLMs)的自动化框架构建而成。该数据集通过三步流水线构建:(1)从阿拉伯语维基百科页面生成带有支持性文本证据的声明;(2)对声明进行变异以生成具有反驳证据的具有挑战性的反事实声明;(3)一个自动验证步骤,用于验证生成的声明是由其附带证据支持还是反驳,或者证据是否不足以判断声明的有效性。由此产生的数据集包含181,976个声明-证据对,标记为支持、反驳或信息不足。对数据集中的测试样本进行的人工评估显示,使用Cohen's Kappa系数,支持声明的标注者间一致性很强(kappa = 0.89),反驳声明的kappa = 0.94。基于人工评估样本的自动验证对支持声明达到了86%的准确率,对反驳声明达到了88%的准确率。为了展示Arafa作为自动阿拉伯语事实核查资源的价值,我们使用Arafa对四个开源基于Transformer的模型进行了微调,其中表现最好的模型在测试数据上达到了77%的宏F1分数。除了Arafa是第一个大规模阿拉伯语事实核查数据集之外,我们的框架还提供了一种可扩展的方法,用于为其他低资源语言开发类似资源。
英文摘要
Automatic fact-checking poses a significant challenge in Arabic natural language processing due to the scarcity of datasets and resources. In this manuscript, we introduce Arafa, a new large-scale dataset for fact-checking in Modern Standard Arabic, constructed through an automated framework leveraging large language models (LLMs). The dataset was constructed through a three-step pipeline: (1) claim generation from Arabic Wikipedia pages with supporting textual evidence, (2) claim mutation to generate challenging counterfactual claims with refuting evidence, and (3) an automatic validation step to validate that the generated claims are either supported or refuted by their accompanying evidence, or if the evidence does not provide enough information to judge the validity of the claims. The resulting dataset comprises 181,976 claim-evidence pairs labeled as supported, refuted, or not enough information. Human evaluation carried out on a test sample from the dataset demonstrated strong inter-annotator agreement (kappa = 0.89) using Cohen's Kappa for supported claims and (kappa = 0.94) for refuted claims. Automatic validation based on a human-evaluated sample achieved 86% accuracy for supported claims and 88% for refuted ones. To showcase Arafa's value as a resource for automatic Arabic fact-checking, four open-source transformer-based models were fine-tuned using Arafa, with the top-performing model achieving a Macro F1-score of 77% on the test data. In addition to Arafa being the first large-scale dataset for Arabic fact-checking, our framework presents a scalable approach for developing similar resources for other low-resource languages.
发表机构
- American University of Beirut(贝鲁特美国大学)
机构由 AI 辅助整理,请以论文原文为准。