arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.00555cs.RO

基于大语言模型的代理探索用于机器人导航与操作的技能编排

LLM-Based Agentic Exploration for Robot Navigation & Manipulation with Skill Orchestration

Abu Hanif Muhammad Syarubany, Farhan Zaki Rahmani, Trio Widianto

更新

AI总结:

本文提出基于大语言模型的代理探索系统,用于机器人导航与操作,通过技能编排实现端到端任务执行。

AI中文摘要:

本文提出了一种端到端的大语言模型(LLM)基于的代理探索系统,用于室内购物任务,在Gazebo仿真和相应的现实世界走廊布局中进行评估。机器人通过在交叉口检测路牌,逐步构建轻量级的语义地图,并存储方向到POI(兴趣点)的关系以及估计的交叉口姿态,同时AprilTags提供可重复的锚点用于接近和对齐。给定一个自然语言的购物请求,LLM在每个交叉口生成一个受约束的离散动作(方向和是否进入商店),并通过ROS有限状态主控制器执行决策,该控制器通过门控模块化运动原语进行执行,包括基于局部成本图的障碍物避障、AprilTag接近、商店进入和抓取。定性结果表明,该集成堆栈可以完成从用户指令到多商店导航和物体检索的端到端任务执行,同时通过其基于文本的地图和记录的决策历史保持模块化和可调试性。

英文摘要:

This paper presents an end-to-end LLM-based agentic exploration system for an indoor shopping task, evaluated in both Gazebo simulation and a corresponding real-world corridor layout. The robot incrementally builds a lightweight semantic map by detecting signboards at junctions and storing direction-to-POI relations together with estimated junction poses, while AprilTags provide repeatable anchors for approach and alignment. Given a natural-language shopping request, an LLM produces a constrained discrete action at each junction (direction and whether to enter a store), and a ROS finite-state main controller executes the decision by gating modular motion primitives, including local-costmap-based obstacle avoidance, AprilTag approaching, store entry, and grasping. Qualitative results show that the integrated stack can perform end-to-end task execution from user instruction to multi-store navigation and object retrieval, while remaining modular and debuggable through its text-based map and logged decision history.

↑