UniWAM技术报告:通过混合流世界-动作建模与操作锚点姿态监督实现统一移动操作
UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision
查看机构详情
- Physical Superintelligence Lab, Fysics AI(Fysics AI 物理超级智能实验室)
- College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
UniWAM通过混合流世界-动作模型和操作锚点姿态监督,实现统一的移动操作,利用大规模自动构建的MAP-Data,在导航和操作任务上显著超越基线。
中文摘要 AI 辅助
移动操作要求精确导航至可操作姿态,随后进行可靠的对象交互。这两个阶段在动作空间和视觉需求上存在差异,这使统一策略学习变得复杂。此外,收集带有明确操作就绪姿态监督的多样化真实世界导航数据成本高昂且难以扩展。我们提出UniWAM,一种统一的混合流世界-动作模型,具有用于导航和操作的独立动作编码器和输出头,并共享一个公共骨干网络。该设计支持在独立采样的导航和操作数据上进行联合表示学习。UniWAM支持任一流的独立推理以及两者的批量并行推理。我们进一步引入操作锚点姿态(MAP)监督,用于确定停止位置和操作朝向。一个自动化流程从大规模3D场景构建MAP-Data,产生超过150万集和7500小时的数据。MAP-Data提供每帧目标对象边界框和图像平面MAP坐标作为辅助导航监督。结合用于操作的投影末端执行器轨迹,这些预测目标为从自我中心观察进行动作学习提供特定流的图像平面监督。借助大规模MAP-Data,UniWAM在MAP-Bench上相比最强外部基线,位置误差降低30.1%,航向误差降低44.0%。在24个真实机器人任务中,UniWAM在MAP导航和移动操作方面取得领先结果,操作性能具有竞争力。我们已发布代码、数据和基准。
英文摘要
Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in action spaces and visual requirements, which complicates unified policy learning. In addition, collecting diverse real-world navigation data with explicit manipulation-ready pose supervision remains costly and difficult to scale. We introduce UniWAM, a unified mixed-stream world-action model with separate action encoders and output heads for navigation and manipulation, sharing a common backbone. This design supports joint representation learning on independently sampled navigation and manipulation data. UniWAM supports independent inference for either stream and batch-parallel inference for both. We further introduce Manipulation Anchor Pose (MAP) supervision for where to stop and how to orient for manipulation. An automated pipeline constructs MAP-Data from large-scale 3D scenes, yielding over 1.5 million episodes and 7,500 hours. MAP-Data provides per-frame target-object bounding boxes and image-plane MAP coordinates as auxiliary navigation supervision. Together with projected end-effector trajectories for manipulation, these prediction targets provide stream-specific image-plane supervision for action learning from egocentric observations. With large-scale MAP-Data, UniWAM outperforms the strongest external baselines on our MAP-Bench by 30.1\% in position error and 44.0\% in heading error. Across 24 real-robot tasks, UniWAM achieves leading results in MAP navigation and mobile manipulation, with competitive manipulation performance. We have released code, data, and benchmark.