arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.04217cs.ROcs.AI

OWMM-Agent:基于多模态智能体数据合成的开放世界移动操作

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

  • Shanghai AI Laboratory(上海人工智能实验室)
  • School of Computing, National University of Singapore(新加坡国立大学计算机学院)
  • Shanghai Jiaotong University(上海交通大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Junting Chen, Haotian Liang, Lingxiao Du, Weiyun Wang, Mengkang Hu, Yao Mu, Wenhai Wang, Jifeng Dai, Ping Luo, Wenqi Shao, Lin Shao

更新

AI总结:

针对开放世界移动操作任务,提出多模态智能体架构与数据合成流程,微调OWMM-VLM实现全局理解、状态跟踪和动作生成,达到SOTA性能并具零样本泛化能力。

AI中文摘要:

导航、操作和视觉模型的快速进步使得移动操作器在许多专门任务中表现出色。然而,开放世界移动操作(OWMM)任务仍然是一个挑战,原因在于需要泛化到开放式指令和环境,以及将高层决策与基于全局场景理解和当前智能体状态的低层机器人控制相集成的系统性复杂性。为了解决这一复杂性,我们提出了一种新颖的多模态智能体架构,该架构维护多视角场景帧和智能体状态以进行决策,并通过函数调用来控制机器人。第二个挑战是领域偏移引起的幻觉。为了增强智能体性能,我们进一步引入了一种用于OWMM任务的智能体数据合成流程,通过指令微调使VLM模型适应我们的任务领域。我们强调,我们微调的OWMM-VLM是首个专为移动操作器设计的专用基础模型,它在统一模型中实现了全局场景理解、机器人状态跟踪和多模态动作生成。通过实验,我们证明了我们的模型在包括GPT-4o在内的其他基础模型中达到了最先进的性能,并在现实世界中展现出强大的零样本泛化能力。项目页面位于https://github.com/HHYHRHY/OWMM-Agent。

英文摘要:

The rapid progress of navigation, manipulation, and vision models has made mobile manipulators capable in many specialized tasks. However, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization to open-ended instructions and environments, as well as the systematic complexity to integrate high-level decision making with low-level robot control based on both global scene understanding and current agent state. To address this complexity, we propose a novel multi-modal agent architecture that maintains multi-view scene frames and agent states for decision-making and controls the robot by function calling. A second challenge is the hallucination from domain shift. To enhance the agent performance, we further introduce an agentic data synthesis pipeline for the OWMM task to adapt the VLM model to our task domain with instruction fine-tuning. We highlight our fine-tuned OWMM-VLM as the first dedicated foundation model for mobile manipulators with global scene understanding, robot state tracking, and multi-modal action generation in a unified model. Through experiments, we demonstrate that our model achieves SOTA performance compared to other foundation models including GPT-4o and strong zero-shot generalization in real world. The project page is at https://github.com/HHYHRHY/OWMM-Agent

补充信息

↑