SuperNav:适用于任意场景任意任务的智能体导航系统
SuperNav: An Agentic Navigation System for Any Task in Any Scene
浏览论文内容
中文总结 AI 辅助
SuperNav为预训练MLLM配备专用智能体 harness,无需导航微调,通过导航技能等组件实现通用导航,在多类任务和环境中优于基线,可部署于真实机器人
中文摘要 AI 辅助
通用服务机器人需要导航系统能够在陌生环境中处理多样化的人类请求,兼顾任务通用性与场景通用性。现有部分方法对多模态大语言模型(MLLM)进行微调以预测导航动作,其行为依赖导航训练数据的覆盖范围,可能限制对新请求和新环境的泛化能力。本文的核心见解是:让MLLM专注于解释请求、理解场景和做出决策,同时保留其通用能力,将运动执行委托给导航工具。为实现这一思路,我们推出SuperNav,它为预训练的MLLM配备了专用智能体 harness,无需对MLLM进行导航特定的微调。该 harness 通过导航技能(面向智能体的物理交互工具)、任务进度与上下文管理来支持这些决策;统一视觉点界面将决策与运动相连,允许模型直接在图像中指定目标,并根据执行反馈修正决策。这些组件共同支持跨不同任务需求和环境的持续导航。SuperNav在实例级、多目标和需求驱动任务上的表现优于4个评估基线;在HM3D上的类别级评估以及在真实四足机器人上的部署,进一步证明了其跨环境适用性。项目页面:this https URL
英文摘要
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments. Our key insight is to let the MLLM focus on interpreting requests, understanding scenes, and making decisions while preserving its general-purpose capabilities and delegating motion execution to navigation tools. To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM. Our harness supports these decisions with Navigation Skills, agent-oriented Tools for physical interaction, and task-progress and context management. A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback. Together, these components support sustained navigation across different task requirements and environments. SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven tasks. Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments. Project Page: https://zju3dv.github.io/SuperNav/
发表机构
- Zhejiang University(浙江大学)
- Shenzhen University(深圳大学)
- Causa Robotics(卡扎机器人公司)
机构由 AI 辅助整理,请以论文原文为准。