AI 中文总结
该研究针对部分可观测场景下LLM智能体的决策挑战,提出NeSyFS神经符号快慢思考框架,结合知识图谱、TSMC算法与反思模块,在三个基准测试中表现优于现有方法。
AI 中文摘要
近年来,大型语言模型(LLMs)越来越多地被部署为自主智能体,应用于自我反思、检索增强生成和科学发现等场景。在这些场景中,智能体必须基于有限的观测而非完整的环境状态采取行动,从而产生部分可观测性,这带来了信念状态推理、任务目标不一致性以及不确定性下规划等关键挑战。现有方法通常基于完整或汇总的动作-观测历史来调整动作,其中冗余和不相关的信息可能误导LLM智能体的决策。受人类认知启发,我们提出了一种新颖的神经符号快慢思考(NeSyFS)框架,用于LLM智能体,以统一的方式解决部分可观测性带来的挑战。我们使用知识图谱(KG)来表示信念状态,为NeSyFS的每个模块提供三元组作为上下文。快思考模块执行反应式动作,而慢思考模块则遵循扭曲序贯蒙特卡洛(TSMC)算法的高层结构进行新的感知不确定性规划。为缓解任务目标不一致性,我们使用反思模块对快思考动作进行反思,并且每当反应式动作反复失败时,就切换到慢思考模块。在ALFWorld、Webshop和ScienceWorld三个代表性基准上的实验表明,该框架相比之前的方法具有显著优势。
英文摘要
Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based on limited observations rather than full environmental states, leading to partial observability. This introduces several key challenges: belief state inference, task objective misalignment, and planning under uncertainty. Prior approaches typically condition actions on full or summarized action-observation histories whose redundant and irrelevant information can mislead the decision making of LLM agent. Inspired by human cognition, we propose a novel neuro-symbolic fast-slow thinking (NeSyFS) framework for LLM agent, addressing the challenges introduced by partial observability in a unified approach. We use a knowledge graph (KG) to represent the belief state, providing triplets as context for every module of NeSyFS. The fast-thinking module performs reactive action, while slow-thinking conducts a new uncertainty-aware planning by following the high-level structure of twisted sequential Monte Carlo (TSMC) algorithm. To mitigate the misalignment of task objective, a reflection module is used to reflect fast-thinking actions, and also switches to the slow-thinking module whenever reactive actions repeatedly fail. Experiments on three representative benchmarks, i.e. ALFWorld, Webshop, and ScienceWorld, demonstrate significant advantages over previous methods.