发表机构
Gaoling School of Artificial Intelligence, Renmin University of China; Beijing Academy of Artificial Intelligence; Tsinghua University; The Hong Kong Polytechnic University(中国人民大学高瓴人工智能学院; 北京人工智能研究院; 清华大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GraphThink框架结合任务图与场景图,通过上下文提示、迭代优化、GRPO奖励设计及事件驱动重规划,在ALFRED基准上实现最优性能,具强泛化能力。
AI 中文摘要
采用基于大语言模型(LLM)规划器的具身智能体常面临物理幻觉、长视距任务泛化能力差、环境感知不足等问题。我们提出GraphThink,这是一种新型框架,它集成任务图以提供结构化知识用于鲁棒规划,同时集成场景图以维护环境记忆用于事件驱动的重规划。具体而言,任务图通过上下文提示和迭代优化引导LLM思维,有效缓解规划幻觉;此外,在GRPO框架内,任务图提供精细的奖励设计以训练LLM规划器,增强长视距规划能力并提升泛化性;最后,由场景图驱动的事件驱动重规划模块实现闭环环境感知与错误修正。GraphThink在ALFRED基准上实现了最先进的性能,尤其在验证集和未见过的长视距任务上,我们的高层规划器优于领先的基于API的LLM,凸显了其强大的零样本和少样本能力;额外评估进一步证明其对新任务和环境具有较强的分布外泛化能力。
英文摘要
Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose GraphThink, a novel framework that integrates a task graph to provide structured knowledge for robust planning and a scene graph to maintain environmental memory for event-driven replanning. Specifically, the task graph guides LLM thinking through contextual prompting and iterative refinement, effectively mitigating planning hallucinations. Furthermore, within the GRPO framework, the task graph offers delicate reward design to train the LLM planner, enhancing long-horizon planning capabilities and improving generalization. Finally, an event-driven replanning module, powered by the scene graph, enables closed-loop environment awareness and error correction. GraphThink achieves state-of-the-art performance on the ALFRED benchmark. In particular, our high-level planner surpasses leading API-based LLMs on both the validation set and held-out long-horizon tasks, underscoring its robust zero-shot and few-shot capabilities. Additional evaluations further demonstrate strong out-of-distribution generalization to novel tasks and environments.