arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你在跟我讲逻辑吗?评估语言模型的三段论推理能力

Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities

Hanna Abi Akl, Fabien Gandon, Catherine Faron, Pierre Monnin

arXiv 2608.12374首次发表:更新:

发表机构

Université Côte d’Azur; Inria; CNRS; I3S; Data ScienceTech Institute(科特阿祖尔大学; 法国国家信息与自动化研究所; 法国国家科学研究中心; 信息科学与系统实验室; 数据科学技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过扩展FOLIO和P-FOLIO数据集,探究不同KR符号对SLMs三段论推理的影响,提出SEF分类法并开源CLGC框架,为提升小型模型推理能力提供了新方法。

AI 中文摘要

语言模型(LMs)在三段论推理等逻辑任务上表现不佳。已有研究表明,知识表示(KR)在表达输入信息以帮助模型解决任务方面起着关键作用。基于这一观察,我们通过扩展FOLIO和P-FOLIO数据集,研究不同形式化KR符号对三段论推理的影响。我们对监督微调(SFT)和零样本(ZS)设置下的小型语言模型(SLMs)开展实验,结果显示输入符号的选择可产生与自然语言相当的性能,同时实现更快的推理速度。我们还提出一种三段论分类方法(SEF),并用其为零样本提示补充逻辑定义,从而提升小型模型的推理能力。我们开源了框架Common Logic Grammar Construction(CLGC),它是首个用于自动生成KR符号下的三段论并定义其SEF类别的Python库。

英文摘要

Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation motivates our study of the impact of different formal KR notations on syllogistic reasoning by extending the FOLIO and P-FOLIO datasets. Our experiments on Small Language Models (SLMs) in Supervised Fine-Tuning (SFT) and Zero-Shot (ZS) settings show that the choice of input notation can yield performances competitive with natural language while enabling faster inference. We also propose a syllogistic categorization method (SEF) and use it to enrich ZS prompts with logical definitions, which boost reasoning in small models. We open-source our framework, Common Logic Grammar Construction (CLGC), as the first Python library for automatically generating syllogisms in KR notations and defining their SEF categories.

CommentsAccepted to the International Joint Conference on Rules and Reasoning (RuleML+RR) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑