MachEmbodied-U0:面向具身智能的统一理解与生成模型
MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence
浏览论文内容
中文总结 AI 辅助
提出统一具身基础模型ME-U0,通过Mixture-of-Transformers架构融合理解与生成,利用子任务预测和可供性接地引导视觉动态与动作生成,在仿真和真实世界任务中取得优异性能。
中文摘要 AI 辅助
通用机器人控制要求模型理解任务意图、识别交互位置、捕捉场景演变方式并生成精确动作。视觉-语言-动作模型提供了强大的语义先验,但通常不显式建模场景动态,而世界-动作模型将视觉预测与控制耦合,却未必暴露细粒度操作所需的任务相关语义和空间结构。我们提出MachEmbodied-U0(ME-U0),一种通过Mixture-of-Transformers架构连接理解与生成专家的统一具身基础模型。子任务预测和可供性接地通过流匹配引导联合视觉动态与动作生成。视觉动态涵盖未来RGB、深度、表面法线和光流,为外观、几何和运动提供互补监督。多速率旋转位置编码(MRPE)将视觉动态与细粒度控制对齐。我们在来自机器人数据集和第一人称视角数据集的约4,200小时精选演示上预训练ME-U0。仅使用每个下游基准中原生可用的监督,ME-U0在RoboDojo仿真基准上取得17.66的平均得分,在LIBERO和LIBERO-Plus上分别取得99.0%和82.5%的平均成功率。我们还在真实世界机器人操作任务上验证ME-U0,证明其在仿真之外的有效性。在没有相应下游监督的情况下,ME-U0进一步在仿真和真实世界观测上展示了零样本子任务预测、可供性接地和视觉动态能力。总体而言,ME-U0在仿真和真实世界中结合了具有竞争力的下游控制性能与可迁移的任务接地和视觉动态能力。
英文摘要
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.
发表机构
- Li Auto Inc(理想汽车)
机构由 AI 辅助整理,请以论文原文为准。