具身导航器:指向、思考、记忆与对齐以实现高效导航
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
- ZJU(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出具身导航框架TAMP-Nav,通过像素转3D动作、选择性推理记忆与GRPO两级对齐,在R2R-CE数据集上实现66.2% SR,仅需90k训练轨迹,达成高效具身导航的最优性能。
AI中文摘要:
尽管大型视觉语言模型(VLMs)已显著推进了具身导航的发展,但其直接部署仍存在挑战——现有方法常迫使VLMs进入与其2D预训练先验不一致的非自然动作空间,加之僵化的推理调度与低效的记忆管理,进一步加剧了问题。为克服这些局限,我们提出TAMP-Nav,一种用于高效具身导航的统一框架。首先,我们引入像素到3D动作形式(指向),将导航重新表述为2D视觉提示:VLMs仅需选择2D像素,这些像素随后被投影为3D坐标以输入低级SLAM控制器,该设计自然使具身执行与VLM固有的2D视觉能力对齐。其次,我们提出集成的选择性推理与锚点-轨迹记忆机制(思考与记忆),该机制在关键节点动态触发思维链,仅保留高保真记忆,将冗余轨迹压缩为轻量型时空指示器,从而保留关键历史信息并增强时空感知。最后,我们通过组相对策略优化(GRPO)设计高效的两级对齐范式(对齐),通过叠加全局结果奖励与细粒度过程奖励,这种密集监督使智能体的认知规划与物理环境反馈紧密对齐,赋予模型自适应推理能力。实验表明,TAMP-Nav实现了最先进的性能(例如在R2R-CE上的成功率SR为66.2%),且具有高运行时与样本效率(仅需90k条训练轨迹)。
英文摘要:
Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).