发表机构
Hanyang University; Hanyang University ERICA(汉阳大学; 汉阳大学ERICA校区)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一个非侵入式框架,通过MCP协议将LLM推理与ROS导航连接,利用视觉地图和语义注释模块,实现超过97%地图覆盖率及基于自然语言指令的空间与语义导航目标选择。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用作机器人系统的自然语言接口,然而它们与基于机器人操作系统(ROS)的导航的集成仍受限于两个差距。首先,诸如占用栅格之类的导航数据以原始几何消息的形式表示,这使得LLMs难以直接将其用作空间或语义上下文。其次,增加LLM驱动的能力通常需要自定义包装器或特定于机器人的接口,限制了跨系统的重用。为了解决这些挑战,我们提出了一个非侵入式框架,通过面向导航的表示层将LLM推理与基于ROS的导航连接起来,该表示层通过模型上下文协议(MCP)作为标准化、可重用的工具暴露,以便任何兼容MCP的LLM都可以访问它们,而无需特定于机器人的包装器。视觉地图模块将占用栅格转换为度量、姿态感知的图像,用于目标推理,而语义注释模块则记录带有机器人姿态的航点级观测。我们在三个任务上评估该框架:自主建图、基于空间推理的导航和基于语义推理的导航。结果表明,所评估的LLM后端使用这些表示在模拟室内环境中实现了超过97%的地图覆盖率,并从自然语言指令中选择空间或语义导航目标。这展示了无需修改现有ROS导航栈即可实现表示介导的LLM导航。
英文摘要
Large language models (LLMs) are increasingly used as natural-language interfaces for robotic systems, yet their integration with Robot Operating System (ROS)-based navigation remains limited by two gaps. First, navigation data such as occupancy grids are represented as raw geometric messages that are difficult for LLMs to use directly as spatial or semantic context. Second, adding LLM-driven capabilities often requires custom wrappers or robot-specific interfaces, limiting reuse across systems. To address these challenges, we propose a non-invasive framework that connects LLM reasoning with ROS-based navigation through a navigation-oriented representation layer, exposed through the Model Context Protocol (MCP) as standardized, reusable tools so that any MCP-compatible LLM can access them without robot-specific wrappers. The visual map modules transform occupancy grids into metric, pose-aware images for goal reasoning, while the semantic annotation modules record waypoint-level observations with robot poses. We evaluate the framework on three tasks: autonomous mapping, spatial reasoning-based navigation, and semantic reasoning-based navigation. The results show that the evaluated LLM backends use these representations to achieve over 97% map coverage and select spatial or semantic navigation targets from natural-language instructions in a simulated indoor environment. This demonstrates representation-mediated LLM navigation without modifying the existing ROS navigation stack.