从语言到导航目标:一种使用RGB-D感知的移动机器人语义导航的视觉-语言方法
From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception
浏览论文内容
中文总结 AI 辅助
研究如何让移动机器人根据自然语言请求进行语义导航,提出基于ROS 2组件的语言驱动导航框架,能识别目标、估计位置并生成导航目标,经仿真和实际场景评估,证明该框架可实现直观人机交互。
中文摘要 AI 辅助
自然语言交互为非专业用户与机器人平台通信提供了直观方式。然而,将用户请求转化为可执行导航行动仍具挑战,需整合语言理解、环境感知和自主导航。本文提出语言驱动导航框架,由模块化ROS 2组件构成,将自然语言指令转化为导航行动。给定自然语言请求,系统识别目标、用RGB-D数据估计其位置并生成导航目标,通过ROS 2 Nav2导航栈执行。在仿真和真实场景中用TurtleBot3 Waffle和Unitree Go2机器人评估,结果表明框架能成功解读指令、生成反馈并导航至目标,证明了结合语义感知和自主导航提供直观人机交互范式的可行性,代码将开源。
英文摘要
Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of language understanding, environment perception, and autonomous navigation. This work presents a language-driven navigation framework that enables mobile robots to interpret user requests in natural language to move the robot to a destination and autonomously navigate towards it. The framework is composed of modular ROS 2 components that cooperate to transform natural language instructions into navigation actions. Given a natural language request referring to a target in the environment (e.g., "go to the mail box"), the system identifies the referenced object, estimates its position using RGB-D data, and generates a navigation goal, which is then executed through the ROS 2 Nav2 navigation stack. The ROS 2-based implementation facilitates portability across different robotic platforms, requiring only the configuration of the corresponding topics and services. The system is evaluated in both simulation and real-world scenarios using a TurtleBot3 Waffle and a Unitree Go2 robot with a RealSense camera. Experimental results show that the framework successfully interprets both direct commands and contextual requests, generates meaningful natural-language feedback, and navigates towards the desired target. These results demonstrate the feasibility of combining semantic perception and autonomous navigation to provide an intuitive human-robot interaction paradigm. Code will be released as open source upon acceptance.