发表机构
Leuphana University of Lüneburg; Mohamed bin Zayed University of Artificial Intelligence(吕讷堡莱普芬纳大学; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对约束密集型逻辑推理任务中思维链提示的不足,提出SymStep方法,通过语言模型分步声明、约束传播器检查一致性等进行推理,在多个逻辑推理任务上表现优异,最小剩余值指导和一致性检查起关键作用。
AI 中文摘要
思维链提示在约束密集型逻辑推理任务中可能严重失败,未经验证的错误会在步骤中悄然累积。我们引入SymStep:一个语言模型每次提出一个原子声明,然后一个轻量级约束传播器检查该声明与先前接受的推导的一致性,拒绝矛盾并自动级联隐含事实。SymStep+G在每个接受的步骤后还提供最小剩余值指导。在ZebraLogicBench的35个谜题保留子集、AR-LSAT分析推理问题、LGP-14等任务上进行实验,结果表明SymStep+G等变体在约束密集型和算术任务上表现出色。消融研究揭示了最小剩余值指导和一致性检查的关键作用。
英文摘要
Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an LLM makes one atomic claim at a time (DEDUCE: Alice, pet, Cat), then a lightweight constraint propagator checks the claim for consistency with prior accepted deductions, rejects contradictions, and cascades implied facts automatically. SymStep+G additionally provides MRV guidance after each accepted step, directing the LLM toward the most constrained unresolved variable. On a 35-puzzle retained subset of ZebraLogicBench, a benchmark of 1,000 Einstein-style logic puzzles, Direct and CoT both achieve 0%, while SymStep+G reaches 97%. On AR-LSAT analytical reasoning problems, SymStep achieves 100% vs. CoT's 87%. On LGP-14, SymStep+G achieves 100% vs. 0% for CoT and Logic-LM, the strongest prior symbolic+LLM baseline we compare against. Ablation studies reveal that MRV guidance is a key mechanism for reducing directionless cycling, while consistency checking provides a safety net against explicit contradictions. Across six benchmarks spanning five task domains, SymStep variants match or exceed every baseline on constraint-dense and arithmetic tasks. Experiments on AQUA-RAT algebra confirm the advantage is constraint-density-specific.