arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34660cs.CL

奖励新颖演绎:求解器引导的过程奖励用于逻辑推理

Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning

Muhammad Asif Ali, Wenqing Wang, Huan Wang, Mohammad Raza

首次发表
浏览论文内容

中文总结 AI 辅助

针对逻辑推理中过程监督不足的问题,提出SPRING方法,利用SMT求解器验证中间步骤并设计过程奖励,鼓励新颖推理,在多个基准上显著提升准确率。

中文摘要 AI 辅助

逻辑推理仍然是大型语言模型(LLMs)面临的主要挑战,尤其是在需要精确约束跟踪、一致性保持和多步演绎的结构化问题上。这一挑战对于小型LLM尤为严峻,因为它们更容易产生不一致、冗余或脆弱的推理轨迹。现有的改进逻辑推理的方法主要优化最终答案的正确性,对中间推理过程仅提供弱监督。在这项工作中,我们提出了SPRING:(求解器引导的新颖逻辑推理步骤生成的过程奖励)。SPRING使用SMT求解器作为训练时中间推理步骤的验证器,以提供过程级监督。它引入了新颖推理步骤的概念,即逻辑上有效、与不断演化的推理状态一致,且未被先前接受的非矛盾演绎所蕴含的步骤。基于这种求解器评估,它设计了过程奖励,鼓励新颖的推理进展,同时惩罚矛盾和无信息的推理步骤。在三个逻辑推理基准(ZebraLogic、AR-LSAT和Knights and Knaves)和四个LLM上的评估显示,SPRING始终优于基础LLM、仅结果奖励基线和Logic-LM。在ZebraLogic上,SPRING相对于基础LLM和最强仅结果基线分别将谜题准确率提高了最多49.71和15.43个百分点。在AR-LSAT上,它分别将整体准确率提高了最多64.93和12.14个百分点。在Knights and Knaves上,SPRING达到了最高93.14的谜题准确率和96.05的人物准确率。

英文摘要

Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING: (Solver-guided Process Rewards for Novel LogIcal ReasoNing Step Generation). SPRING uses SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. It introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, it designs process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation across three logical reasoning benchmarks, ZebraLogic, AR-LSAT, and Knights and Knaves, and four LLMs shows that SPRING consistently outperforms base LLMs, outcome-only reward baselines, and Logic-LM. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and strongest outcome-only baseline, respectively. On AR-LSAT, it improves overall accuracy by up to 64.93 and 12.14 points, respectively. On Knights and Knaves, SPRING achieves up to 93.14 puzzle accuracy and 96.05 person accuracy.

发表机构

  • Information Technology University(信息技术大学)
  • Huazhong Agricultural University(华中农业大学)
  • Qatar Computing Research Institute, Hamad Bin Khalifa University(卡塔尔计算研究所,哈马德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑