arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32491cs.CLcs.AI

用词汇蕴含解释文本蕴含:利用LLM为形式证明提供词汇关系

Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs

Jorryt de Jong, Stefan Moraca, Ettore Cesari, Lasha Abzianidze

首次发表
浏览论文内容

中文总结 AI 辅助

本文评估LLM在逻辑NLI系统中提供词汇蕴含的能力,发现其生成关系部分合理且贡献有限,任务仍具挑战性。

中文摘要 AI 辅助

大语言模型(LLMs)在自然语言推理方面能力很强,并且似乎存储了大量词汇知识,但目前仍不清楚它们在推理时实际使用了多少这些知识,以及是否以正确的方式使用。另一方面,基于逻辑的自然语言推理(NLI)系统提供了透明且形式化的推理基础,但需要提供丰富的词汇知识来证明超出纯逻辑推理的推断。在本文中,我们评估LLM能否识别解决NLI问题所需的全部词汇知识,以及这些知识在基于逻辑的NLI系统中对证明搜索的贡献程度。我们的研究仅聚焦于结构化词汇蕴含(例如,chinchilla⊑small animal)作为带有蕴含标签的NLI问题的结构化解释的代理。首先,我们为一个新任务整理了一个数据集,该任务是用一组词汇蕴含来解释句子蕴含。该数据集用于对LLM生成结构化词汇解释进行内在评估。然后,我们在一个简单的神经符号设置中使用NLI作为外在评估,评估LLM能否为自然语言的自然逻辑定理证明器LangPro提供足够的词汇关系。结果表明,即使对托管专有LLM而言,所提出的任务仍然具有挑战性,并且它们对定理证明的贡献是适度的:生成的关系通常仅部分合理,可能针对特定NLI问题定制,而非代表普遍有效的词汇知识。

英文摘要

Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. On the other hand, logic-based Natural Language Inference (NLI) systems provide transparent and formally grounded reasoning, but they need to be supplied with rich lexical knowledge to prove inferences beyond purely logical ones. In this paper, we evaluate whether LLMs can identify all lexical knowledge needed to solve NLI problems and how much this knowledge contributes to proof search in a logic-based NLI system. Our research focuses exclusively on structured lexical entailments (e.g., chinchilla$\sqsubseteq$small animal) as a proxy for structured explanations for NLI problems with an entailment label. First, we curate a dataset for a new task of explaining sentential entailments with a set of lexical entailments. The dataset is used to intrinsically evaluate LLMs on generating structured lexical explanations. Then, we use NLI as an extrinsic evaluation in a simple neuro-symbolic setting, assessing whether LLMs can supply sufficient lexical relations to LangPro, a natural-logic theorem prover for natural language. The results show that the proposed task remains challenging even for hosted proprietary LLMs, and that their contribution to theorem proving is moderate: generated relations are often only partially sound and may be tailored to the specific NLI problem rather than representing generally valid lexical knowledge.

补充信息

↑