arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

E-CONAN(蕴含、矛盾与中立)诊断数据集:探究阿拉伯语自然语言理解中的语言现象

E-CONAN (Entailment, CONtradition And Neutral) Diagnostics Dataset Investigating Linguistic Phenomena in Arabic Natural Language Understanding

Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi

arXiv 2609.33530首次发表:更新:

发表机构

Higher Institute for Applied Sciences and Technology; Arab International University(高等应用科学与技术学院; 阿拉伯国际大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对阿拉伯语NLU,提出E-CONAN诊断数据集和层级错误分析框架,评估9个预训练模型和5个LLM,发现LLM在知识和常识推理上更强,但在句法上较弱。

AI 中文摘要

自然语言理解(NLU)在各种应用中扮演着关键角色,但其性能在处理人类语言的复杂性方面存在不足,这些问题涵盖从词汇歧义到高级推理困难。跨多种语言现象的错误分析对于改进NLU至关重要,因为这有助于人类深入理解模型的局限性和能力,从而优化模型的泛化能力。值得注意的是,多个基准测试包含用于调查和细粒度错误分析的诊断数据集。在指出当前最先进技术的空白时,我们注意到宏观和微观类别没有命名约定,甚至没有一套应涵盖的标准语言现象。为弥补这一空白,我们提出了一个跨语言NLU错误分析的初始层级结构。此外,我们提出了一种创建NLI层级框架的方法,并将其应用于阿拉伯语NLU的案例研究。此外,本文介绍了E-CONAN诊断数据集,这是一个免费提供的数据集,基于我们提出的阿拉伯语层级结构,对粗粒度和细粒度类别进行了人工标注。E-CONAN数据集通过错误分析和深入调查帮助NLU设计者更好地理解其模型。我们使用E-CONAN调查了9个预训练语言模型和5个LLM的性能。结果表明,LLM在世界知识和常识推理宏观类别上优于预训练模型,而在句法宏观类别上不如预训练模型。此外,所有模型最难的现象是推理,所有预训练模型最容易的现象是句法,而LLM最容易的是词汇-句法。

英文摘要

Natural Language Understanding (NLU) plays a crucial role in various applications, yet its performance suffers from weaknesses in handling the complexities of human languages, ranging from lexical ambiguity to high-level reasoning difficulties. Analyzing errors across diverse linguistic phenomena is crucial for NLU improvement, as it will help humans get insights to comprehensively assess models' limitations and capabilities, so optimizing models' generalization. Notably, several benchmarks contain diagnostics datasets designed for investigation and fine-grained error analysis. When highlighting the gaps in the state-of-the-art, we noted that there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena that should be covered. To overcome this gap, we propose an initial hierarchy for Cross-Lingual NLU error analysis. Moreover, we propose a methodology to create an NLI hierarchical framework and applied a case study on Arabic NLU. Moreover, this paper introduces E-CONAN diagnostics dataset, a freely available dataset manually-annotated with coarse-grained and fine-grained categories based on our proposed Arabic hierarchy. E-CONAN dataset helps NLU designers better understand their models by doing error analysis and in-depth investigation. We used E-CONAN to investigate the performance of 9 pretrained language models and 5 LLMs. Results indicate that LLMs outperform pretrained models in world knowledge and commonsense reasoning macro-category, and underperform pretrained models in syntactic macro-category. Moreover, the hardest phenomena for all models is Reasoning, and the easiest phenomena for all pretrained models is Syntactic, and the easiest for LLMs is Lexico-Syntactic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑