arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小米机器人-U0:基于世界基础模型的统一具身合成

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li

arXiv 2607.11643首次发表:更新:

发表机构

Xiaomi Robotics(小米机器人)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对基础模型应用于具身场景受限问题,提出小米机器人-U0多模态自回归模型,将具身生成扩展为基础图像和视频生成的形式,联合优化多种生成任务,在多方面取得领先成果,证明基础世界模型可用于具身智能。

AI 中文摘要

近期的基础图像和视频生成模型具有强大的泛化和可控性,但直接应用于具身场景时受多视图一致性、几何连贯性和机器人实体约束等限制。现有方法通常用有限机器人数据适配基础模型,常牺牲大规模预训练获得的视觉知识。我们提出小米机器人-U0,一个380亿参数的多模态自回归模型用于统一具身合成。它将具身生成视为基础图像和视频生成的扩展,联合优化文本到图像生成、图像编辑、具身场景生成、具身转移和具身视频生成。该统一框架在适配具身设置时保留了预训练世界基础模型的泛化性。小米机器人-U0是首个支持跨多个机器人实体进行高质量多视图场景生成的模型,还引入结构化、可控的具身转移用于细粒度编辑,同时保持多视图一致性和交互动态。它在单步和序列生成任务上取得了领先成果,在具身场景生成和转移的人工评估中优于GPT-Image-2.0,在具身视频生成的世界竞技场中排名第一,在具有挑战性的真实世界操作任务中将pi_0.5的分布外成功率从36.9%提高到63.2%。这些结果表明基础世界模型可作为具身世界模型和具身智能的可扩展数据引擎。代码和检查点可在指定网址获取。

英文摘要

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑