arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PreAct-Nav:城市导航中行动前的智能体推理

PreAct-Nav: Agentic Reasoning Before Action for Urban Navigation

Jing Xie, Shouwei Ruan, Yubin Wang, Yuxiang Zhang, Junwei Yang, Songchang Jin, Dianxi Shi

arXiv 2610.04916首次发表:更新:

发表机构

Shanghai Jiao Tong University; Beihang University; The Hong Kong University of Science and Technology; Tsinghua University; University of Cambridge; Intelligent Game and Decision Lab (IGDL)(上海交通大学; 北京航空航天大学; 香港科技大学; 清华大学; 剑桥大学; 智能博弈与决策实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PreAct-Nav提出了一种智能体导航框架,通过记忆模块和预测性世界沙盒,在行动前进行推理,将远期目标转化为中期子目标,从而提升城市导航中的行动选择,尤其在长路线和转弯多的路线上效果显著。

AI 中文摘要

城市导航要求具身智能体基于自我中心的观察,通过局部决策来追求长期目标。然而,现有的智能体导航方法在大规模物理环境中,往往难以将遥远的目标转化为连贯的局部决策。它们依赖于对瞬时观察或有限历史进行语言推理,这限制了对行动后果和未来状况的预见能力,尽管这种预见对于在漫长而复杂的城市路线中导航至关重要。为了弥合这一差距,我们提出了PreAct-Nav,一个智能体导航框架,为冻结的策略配备预见性推理,以实现稳健的城市导航。我们的核心思想是将局部决策锚定在持久的、中期范围的子目标上,在执行前评估预测行动的后果,并利用实际结果持续更新推理上下文。其核心是一个导航记忆模块,该模块在决策过程中维护活跃子目标和相关经验,将遥远的目标转化为可操作的中间目标。我们进一步引入了一个预测性世界沙盒,使用行动条件世界模型(AC-WM)来预测基于候选移动的世界动态。一个视觉语言模型(VLM)推理器在当前子目标下解释这些预测,以保留或修改行动。执行后,使用真实观察来评估结果,纠正不一致的假设,并更新记忆以继续或重新制定子目标。广泛的评估表明,所提出的PreAct-Nav通过记忆更新和视觉预测改善了行动选择,在更长的路线和转弯更多的路线上收益更为显著。

英文摘要

Urban navigation requires embodied agents to pursue long-horizon goals through local decisions based on egocentric observations. However, existing agentic navigation methods often struggle to translate distant goals into coherent local decisions in large-scale physical environments. Their reliance on linguistic reasoning over transient observations or limited history constrains anticipation of the consequences of actions and future conditions, despite the importance of such foresight for navigating long and complex urban routes. To bridge this gap, we propose PreAct-Nav, an agentic navigation framework that equips frozen policies with anticipatory reasoning for robust urban navigation. Our central idea is to anchor local decisions in persistent medium-horizon subgoals, assess the consequences of predicted actions before execution, and continually update the reasoning context using actual outcomes. At its core, a navigation memory module maintains the active subgoal and relevant experience across decisions, translating distant goals into actionable intermediate objectives. We further introduce a predictive world sandbox that uses an action-conditioned world model (AC-WM) to forecast world dynamics conditioned on candidate movements. A vision-language model (VLM) reasoner interprets these predictions under the current subgoal to retain or revise actions. After execution, real observations are used to assess outcomes, correct inconsistent assumptions, and update memory to continue or reformulate the subgoal. Extensive evaluations demonstrate that the proposed PreAct-Nav improves action selection through memory updates and visual prediction, with more pronounced gains on longer routes and routes with more turns.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑