arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.06537cs.RO

UniLM-Nav:零样本最后一英里导航的统一框架

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation

Zhuofan Zhang, Tianxu Wang, Guoxi Zhang, Yixiong Lin, Xilin Wang, Hongming Xu, Qing Li, Song-Chun Zhu, Lifeng Fan

首次发表
浏览论文内容

中文总结 AI 辅助

研究移动操作中最后一英里导航问题,提出UniLM-Nav统一框架,通过多模态大语言模型后端分解任务为视图选择、功能接地和姿态推理,在OVMM基准上优于现有方法,还验证了在实际机器人上的适用性。

中文摘要 AI 辅助

移动操作要求机器人导航到目标物体或容器并执行预期操作。然而,到达目标附近并不保证有可进行操作的基础姿态,即所谓的最后一英里导航问题。先前方法依赖手动姿态标注或特定任务训练,限制了其在具有细粒度空间约束的开放词汇设置中的扩展性。我们提出UniLM-Nav,一个用于零样本开放词汇最后一英里导航的统一框架。它将最后一英里导航分解为视图选择、任务条件下的功能接地和几何感知基础姿态推理,通过共享的多模态大语言模型后端解决。具体步骤包括首先选择最佳捕捉目标的参考视图,然后在所选视图中确定与任务相关的功能点并转换到机器人中心坐标系,最后根据接地功能、任务上下文和机器人几何形状推断出机器人可操作的基础姿态。我们在OVMM基准上评估了UniLM-Nav,它比之前的最先进方法MoTo高出3.13个百分点。分析表明我们方法的组件对最终性能至关重要,大语言模型的选择也有很大影响。我们还将UniLM-Nav部署在带有6自由度Unitree Z1机械手的Unitree B2四足机器人上,验证了其在实际移动操作任务中的适用性。

英文摘要

Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity of the target does not guarantee a manipulation-ready base pose, a problem known as last-mile navigation. Prior methods for last-mile navigation either rely on manual pose annotation or task-specific training, limiting their scalability to open-vocabulary settings with fine-grained spatial constraints. We propose UniLM-Nav, a unified framework for zero-shot open-vocabulary last-mile navigation. UniLM-Nav decomposes last-mile navigation into view selection, task-conditioned affordance grounding, and geometry-aware base-pose reasoning, all resolved with a shared multimodal large language model (MLLM) backend. Specifically, UniLM-Nav first selects a reference view that best captures the target object or receptacle from recently collected observations. It then grounds task-relevant affordance point in the selected view and lifts the result into the robot-centric coordinate frame. Finally, conditioned on the grounded affordance, task context, and robot geometry, it infers a manipulation-ready base pose for the robot. We evaluate UniLM-Nav on the OVMM benchmark, where it outperforms the previous state-of-the-art method, MoTo, by 3.13 percentage points. Analyses show that the components of our method are crucial to final performance, and that the choice of MLLM also has a substantial effect. We further deploy UniLM-Nav on a Unitree B2 quadruped robot with a 6-DoF Unitree Z1 manipulator, validating its applicability to real-world mobile manipulation tasks.

发表机构

  • Tsinghua University(清华大学)
  • State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司人工智能研究院)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑