arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小米机器人1:利用超过10万小时的真实世界轨迹扩展视觉语言动作模型

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou

arXiv 2607.15330首次发表:更新:

发表机构

Xiaomi Robotics(小米机器人)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出小米机器人1这一视觉语言动作模型,采用两阶段训练方法,预训练利用大量真实轨迹赋予模型能力,后训练使其与机器人实体及指令对齐。该模型在多基准测试中表现优异,能有效微调,建立新的最先进水平。

AI 中文摘要

我们展示了小米机器人1,这是一个基础的视觉语言动作(VLA)模型,能够:(1)遵循各种语言指令,在未见环境中开箱即用执行广泛的移动操作任务;(2)使用最少的微调数据有效适应新的下游任务。我们提出了一个由预训练和后训练组成的两阶段训练方法。预训练期间,通过在超过10万小时通过UMI设备收集的真实世界操作轨迹上训练,赋予模型广泛且可泛化的动作生成能力。关键的是,我们开发了一个可扩展的自动标注管道,用描述场景状态转换的自然语言标注轨迹片段,为动作学习提供丰富而精确的条件。后训练旨在使这些能力与机器人实体以及人类自然用于提示机器人的命令式指令对齐。大量实验证明了强大的扩展行为。小米机器人1在预训练期间随着数据规模和模型大小的增加持续改进。这种扩展行为直接转移到后训练中,更强的预训练模型在未见环境中产生更好的开箱即用真实机器人性能。此外,小米机器人1作为一个强大的机器人基础策略,可以在复杂、灵巧的任务上以高数据效率进行有效微调。在多个模拟基准测试中,小米机器人1优于现有方法。值得注意的是,它在RoboCasa365上以57.6%的成功率建立了新的最先进水平,超过了之前的最佳成绩46.6%。此外,它在RoboDojo上的平均得分为20.07,显著优于先前的最先进水平(13.07)。代码和模型检查点将被发布。项目页面:这个https链接

英文摘要

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.4% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html

CommentsProject page: https://robotics.xiaomi.com/xiaomi-robotics-1.html

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑