arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.17462cs.ROcs.AIcs.CV

通用机器人导航:基于LVLM编排的感知、推理与行动

General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting

  • Stanford University(斯坦福大学)
  • Jet Propulsion Laboratory(喷气推进实验室)
  • California Institute of Technology(加州理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Bernard Lange, Anil Yildiz, Mansur Arief, Shehryar Khattak, Mykel Kochenderfer, Georgios Georgakis

更新

AI总结:

针对未知环境通用导航难题,提出ARNA框架,利用LVLM智能体编排感知、推理与导航工具,实现自主工作流,在HM-EQA基准上超越现有方法,并展现出跨任务泛化能力。

AI中文摘要:

为未知环境开发通用导航策略仍是机器人学的核心挑战。现有系统大多依赖任务特定的神经网络和固定的信息流,限制了其泛化能力。大型视觉-语言模型(LVLMs)通过嵌入类人知识进行推理和规划,提供了一种有前景的替代方案,但先前LVLM-机器人集成主要依赖预建地图、硬编码表示和刚性控制逻辑。我们提出了智能体机器人导航架构(ARNA),这是一个通用框架,为基于LVLM的智能体配备了一组来自现代机器人技术栈的感知、推理和导航工具库。在运行时,智能体自主定义并执行任务特定的工作流,迭代查询模块、对多模态输入进行推理,并选择导航动作。这种智能体式公式化方法使得在先前未映射的环境中实现稳健的导航和推理成为可能,为机器人技术栈设计提供了新视角。在Habitat Lab的HM-EQA基准上评估,ARNA优于最先进的EQA特定方法。在RxR和自定义任务上的定性结果进一步证明了其在广泛导航挑战中的泛化能力。

英文摘要:

Developing general-purpose navigation policies for unknown environments remains a core challenge in robotics. Most existing systems rely on task-specific neural networks and fixed information flows, limiting their generalizability. Large Vision-Language Models (LVLMs) offer a promising alternative by embedding human-like knowledge for reasoning and planning, but prior LVLM-robot integrations have largely depended on pre-mapped spaces, hard-coded representations, and rigid control logic. We introduce the Agentic Robotic Navigation Architecture (ARNA), a general-purpose framework that equips an LVLM-based agent with a library of perception, reasoning, and navigation tools drawn from modern robotic stacks. At runtime, the agent autonomously defines and executes task-specific workflows that iteratively query modules, reason over multimodal inputs, and select navigation actions. This agentic formulation enables robust navigation and reasoning in previously unmapped environments, offering a new perspective on robotic stack design. Evaluated in Habitat Lab on the HM-EQA benchmark, ARNA outperforms state-of-the-art EQA-specific approaches. Qualitative results on RxR and custom tasks further demonstrate its ability to generalize across a broad range of navigation challenges.

↑