arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.28144cs.AI

解构空间复杂性:用于LLM空间推理的层次分解

Deconstructing Spatial Complexity: Hierarchical Decomposition for LLM Spatial Reasoning

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

Yi Wang, Haojie Lu, Zhaofan Zhang, Li Chen, Sihong Xie

更新

AI总结:

提出一种层次任务分解方法,结合MCTS引导的组相对策略优化(M-GRPO),通过改进中间状态选择和规划能力,显著提升LLM在导航、规划和策略游戏等空间任务中的表现。

AI中文摘要:

LLMs在通用语言理解和推理方面表现出色。然而,它们在空间推理方面始终表现不佳,这严重限制了它们的应用,特别是在具身智能领域。受层次强化学习成功的启发,本文介绍了一种新颖的LLM空间推理层次任务分解方法。我们的方法通过识别关键中间状态并生成简化的子环境,引导LLMs将复杂任务分解为可管理的子任务。然而,我们发现LLMs由于缺乏足够的空间先验知识,往往无法推导出最优的中间状态,导致次优的任务分解。为了解决这一限制并增强其规划能力,我们提出了MCTS引导的组相对策略优化(M-GRPO),其中我们通过结合LLM的先验预测概率及其认知不确定性来重新制定UCT公式。此外,我们实现了一个更细粒度的优势函数,使模型能够学习最优路径规划。实验结果表明,我们的方法显著提高了LLM在空间任务(包括导航、规划和策略游戏)上的性能,达到了最先进的结果。这项工作为LLM在现实世界中的应用铺平了道路。

英文摘要:

LLMs have shown remarkable proficiency in general language understanding and reasoning. However, they consistently underperform in spatial reasoning that severely limits their application, particularly in embodied intelligence. Inspired by the success of hierarchical reinforcement learning, this paper introduces a novel method for hierarchical task decomposition in LLM spatial reasoning. Our approach guides LLMs to decompose complex tasks into manageable sub-tasks by identifying key intermediate states and generating simplified sub-environments. However, we identify that LLMs often fail to derive optimal intermediate states due to their insufficient spatial prior, leading to sub-optimal task decomposition. To address this limitation and enhance its planning capability, we propose the MCTS-Guided Group Relative Policy Optimization (M-GRPO), where we reformulate the UCT formula by incorporating the LLM's prior predictive probabilities alongside its epistemic uncertainty. Furthermore, we implement a more fine-grained advantage function, enabling the model to learn optimal path planning. Experimental results demonstrate that our method substantially improves LLM performance on spatial tasks, including navigation, planning, and strategic games, achieving state-of-the-art results. This work paves the way for LLMs in real-world applications.

补充信息

↑