arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

E-CONAN(蕴含、矛盾与中立)基准:阿拉伯语文本蕴含与自然推理数据集

E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi

arXiv 2609.11334首次发表:更新:

发表机构

Higher Institute for Applied Sciences and Technology; Arab International University(高等应用科学与技术学院; 阿拉伯国际大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对阿拉伯语NLI资源匮乏,构建E-CONAN基准数据集(含二分类和三分类),评估多语言预训练模型与LLM,证明其多样构成优于现有基准,助力泛化评估与微调。

AI 中文摘要

自然语言推理处理句子对以提取其语义关系。NLI一直是热门研究课题,并作为其他NLP应用的主要组成部分。尽管世界范围内各种语言的文本推理取得了显著进展,但阿拉伯语在此领域仍面临资源有限的问题。为解决这一差距,本文介绍了E-CONAN基准,其由来自不同来源的句子对组成:(1)自动翻译的句子对,(2)人工验证的机器翻译句子对,(3)从对外阿拉伯语教学书籍中手工构建的句子对,以及(4)来自包含谣言的不同新闻频道的标题对。E-CONAN包含两个基准数据集:E-CONAN-2,一个二分类数据集(RTE)和E-CONAN-3,一个三分类数据集(NLI)。此外,我们使用E-CONAN基准通过零样本分类评估了9个最先进的多语言预训练模型。模型在ArNLI、XNLI和E-CONAN数据集上进行了评估。结果表明,E-CONAN是评估模型泛化能力甚至微调预训练模型的潜在宝贵资源。其多样化的构成,源于多种来源的组合,与XNLI和ArNLI相比,提供了更广泛和更稳健的评估。此外,我们在E-CONAN-3数据集上评估了5个大型语言模型(LLM)。而且,我们纳入了MARBERT作为阿拉伯语特定基线的代表,并进行了性能评估比较,以展示阿拉伯语特定模型在E-CONAN基准上如何与跨语言和基于LLM的方法相抗衡。此外,我们进行了详细的定性和定量错误分析,以分析常见的错误模式。E-CONAN基准将公开提供,我们希望它将丰富阿拉伯语文本蕴含和自然语言推理的研究社区。

英文摘要

Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main component in other NLP applications. Despite significant advancements in textual inference across various languages all around the world, Arabic language still suffers from limited resources in this domain. To address this gap, this paper introduces E-CONAN benchmarks that are composed of sentences pairs from various sources: (1) automatically-translated pairs, (2) human-validated machine-translated pairs, (3) hand-crafted pairs from teaching Arabic as foreign language books, and (4) headlines pairs from different news channels containing rumors. E-CONAN contains two benchmark datasets, E-CONAN-2, a 2-way dataset (RTE) and E-CONAN-3, a 3-way dataset (NLI). Additionally, we have used E-CONAN benchmarks to evaluate 9 state-of-the-art multilingual pretrained models using zero-shot classification. Models were evaluated across the ArNLI, XNLI, and E-CONAN datasets. Results show that E-CONAN is a potentially valuable resource for evaluating model generalization and even for fine-tuning pre-trained models. Its diverse composition, derived from a combination of sources, offers a broader and more robust assessment compared to XNLI and ArNLI. In addition, we have evaluated 5 LLMs on E-CONAN-3 dataset. Moreover, we incorporated MARBERT as a representative Arabic-specific baseline and conducted performance evaluation comparison to demonstrate how Arabic-specific models scale against cross-lingual and LLM-based approaches on the E-CONAN benchmarks. Furthermore, we conducted detailed qualitative and quantitative error analysis to analyze frequent error patterns. E-CONAN benchmarks will be publicly available, we hope that it will enrich research community in Arabic textual entailment and natural language inference.

DOI:10.1109/ACCESS.2026.3732060

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑