arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合理地运用波普尔方法:用归纳逻辑编程修补上下文推理

Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming

Zirong Chen, Meiyi Ma

arXiv 2607.23019首次发表:更新:

发表机构

Vanderbilt University(范德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对思维链提示中间步骤逻辑不合理问题,提出Reason Popper-ly框架,用归纳逻辑编程学习规则作在线验证器校正步骤,在CLUTRR基准测试中提升了模型准确率,还产生细粒度错误分类法,优于完全外部符号管道。

AI 中文摘要

思维链(CoT)提示使大语言模型(LLMs)能够处理多步推理任务,但生成的中间步骤不一定逻辑合理。我们提出了Reason Popper-ly,这是一个神经符号框架,它使用归纳逻辑编程(ILP)从推理轨迹中学习关系合成规则,并将其作为在线验证器进行步骤级校正。给定LLM生成的轨迹,该方法根据学习到的规则表检查每个推理步骤,诊断违规类型,用符号推导的修复重写错误步骤,并重新生成其余后缀,以便模型能够根据经过验证的轨迹产生最终答案。我们在CLUTRR(一个多跳亲属关系推理基准)上进行评估,使用五个语言模型处理2到10跳的推理链。在所有模型中,Reason Popper-ly始终比标准CoT提高最终准确率,对于小模型,在最长链上提高多达48个百分点,对于前沿模型提高15个百分点。与完全外部的符号管道相比,我们的方法在更难的实例上表现更好,通过保留模型的成功基础,同时只纠正可验证的推理失败。此外,步骤级ILP验证产生了一个细粒度的错误分类法,提供了超出最终答案准确率的诊断见解。

英文摘要

Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound. We present Reason Popper-ly, a neurosymbolic framework that uses inductive logic programming (ILP) to learn relation composition rules from reasoning traces and deploys them as an online verifier for step-level correction. Given an LLM-generated trace, the method checks each inferred step against the learned rule table, diagnoses the violation type, rewrites incorrect steps with symbolically derived repairs, and regenerates the remaining suffix so that the model can produce its final answer conditioned on a verified trace. We evaluate on CLUTRR, a multi-hop kinship reasoning benchmark, using five language models over reasoning chains of 2 to 10 hops. Across all models, Reason Popper-ly consistently improves terminal accuracy over standard CoT, with gains of up to 48 percentage points for small models and 15 points for frontier models on the longest chains. Compared with a fully exogenous symbolic pipeline, our method performs better on harder instances by preserving the model's successful grounding while correcting only verifiable reasoning failures. In addition, step-level ILP verification yields a fine-grained error taxonomy that provides diagnostic insight beyond final-answer accuracy.

CommentsAccepted at the 20th Conference on Neurosymbolic Learning and Reasoning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑