发表机构
Xspark AI; The Hong Kong University of Science and Technology (Guangzhou); Tsinghua University; The University of Hong Kong(Xspark AI; 香港科技大学(广州); 清华大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MM-ABC基础模型,通过观察、协调与想象机制,结合稀疏多级VLM特征、未来监督和MM-APT协调模块,在多个基准和真实任务上显著提升移动操作成功率。
AI 中文摘要
移动操作通过使可达区域本身可控,将机器人交互扩展到固定运动学工作空间之外。这种灵活性带来了两个核心挑战:在连续自我运动下的空间感知,以及对异构机械臂与底盘动作的协调控制。现有方法通过显式3D表示或预测性世界模型来强化几何,并常将移动与操作解耦为独立的动作流。我们认为,有效的移动操作不仅需要解耦,还需要支持跨流高效协作的表示。我们提出MM-ABC,一个围绕“观察、协调与想象”臂-基座协作构建的基础模型。MM-ABC结合稀疏多级VLM特征用于空间感知;一个仅用于训练的未来分支,利用世界想象和几何意图作为额外监督,强化感知与操作意图预测,并改善整体学习信号;以及MM-APT,通过掩码联合注意力和干净动作x预测来协调独立的操作与移动流。在受控消融实验中,将干净动作预测替换为速度预测会使RoboCasa365组合可见任务的成功率从32.8%降至29.2%,而移除未来监督或多级条件会导致更大下降。我们在5000多小时的异构机器人数据上预训练MM-ABC,涵盖40万多个片段、12个数据集和17种实体。实验覆盖EBench、RoboCasa365、ManiSkill-HAB、LIBERO、LIBERO-Plus以及真实世界移动操作。MM-ABC在EBench上达到44.71%的成功率,在RoboCasa365上为61.2%,在LIBERO上为99.1%,在无扰动训练的LIBERO-Plus上为82.8%,并在五个真实世界任务上达到83%的平均成功率。
英文摘要
Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.
CommentsWebpage: https://mm-abc.github.io/