发表机构
Georgetown University; Syntheos, Corp; UNU-MERIT(乔治城大学; Syntheos公司; 联合国大学-马斯特里赫特经济与社会创新研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在Clue棋盘游戏环境中,通过工具增强的可能性矩阵将推理日志转为显式状态,提升LLM智能体的多步演绎推理能力,并显著优于基线。
AI 中文摘要
尽管大型语言模型(LLM)近期取得了进展,但在扩展交互中执行逻辑一致的演绎推理仍然具有挑战性。需要跨多个推理步骤整合证据、与先前推断保持一致性以及在新的约束下更新信念的任务,可以暴露当前模型的局限性,同时为评估推理增强提供一个有用的测试平台。在本文中,我们实现了一个基于文本的多智能体版本的经典棋盘游戏Clue,作为评估多步骤、智能体式演绎推理的环境。在此设置中,智能体必须从一系列观察中推断隐藏信息,在回合间保持一致性,并在不断演变的逻辑约束集上进行推理。我们实例化了六个基于LLM的智能体(GPT-4o-mini和Gemini-2.5-Flash)作为参与回合制游戏的玩家;使用每个模型家族三个智能体,我们在重复游戏中建立了基线性能。然后,我们引入了一种工具增强方法,其中结构化的可能性矩阵将生成的推理日志中的隐式游戏状态转换为剩余可能性的显式表示。可能性矩阵编码了扩展回合记忆和演绎约束,将这些任务从智能体中卸载。我们将此方法与基线进行比较,以评估工具增强如何支持战略推理环境中自主智能体的推理质量和任务成功率。
英文摘要
Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging. Tasks that require integrating evidence across multiple reasoning steps, maintaining consistency with prior inferences, and updating beliefs under new constraints can surface limitations in current models while providing a useful testbed for evaluating reasoning enhancements. In this paper, we implement a text-based, multi-agent version of the classic board game Clue as an environment to evaluate multi-step, agentic deductive reasoning. In this setting, agents must infer hidden information from a sequence of observations, maintain consistency across turns, and reason over an evolving set of logical constraints. We instantiate six LLM-based agents (GPT-4o-mini and Gemini-2.5-Flash) as players that engage in turn-based gameplay; using three agents per model family, we establish baseline performance across repeated games. We then introduce a tool-augmented approach in which a structured possibility matrix converts implicit game state from generated reasoning logs into an explicit representation of remaining possibilities. The possibility matrix encodes extended-turn memory and deductive constraints, offloading these tasks from the agent. We compare this approach against the baseline to evaluate how tool augmentation supports reasoning quality and task success for autonomous agents in a strategic reasoning environment.
Journal refIJCAI 2026 Workshop Logical and Symbolic Reasoning