arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Being-M0.7:一种用于人形机器人的潜在世界-动作模型

Being-M0.7: A Latent World-Action Model for Humanoid Robots

Junpeng Yue, Boyuan Li, Yuxuan Wang, Zepeng Wang, Yuhui Fu, Feiyang Xie, Yu Zhang, Jing Zhang, Xianqi Zhang, Weibo Li, Xiaofei Zheng, Yuming Fang, Jiangxing Wang, Zongqing Lu

arXiv 2610.11283首次发表:更新:

发表机构

BeingBeyond(BeingBeyond)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Being-M0.7是一种人形机器人潜在世界-动作模型,通过多模态人类数据预训练等三阶段训练,在SIMPLE基准获最高综合成功率,在Unitree G1真实任务上与最强基线表现相当。

AI 中文摘要

人形机器人的移动操作需要协调的 locomotion(移动)与 manipulation(操作),且需依据未来场景演化和全身运动来指导,但学习这些能力受限于稀缺的机器人演示数据。人类视频和运动数据集可提供可扩展的监督信号,但许多仅包含视频或运动数据,而非视频-运动配对数据;此外,人类运动无法直接指定机器人可执行的动作。我们提出 Being-M0.7,一种潜在世界-动作模型,通过预训练、机器人中期训练和动作后期训练,将从多模态人类数据中学习到的视觉-运动先验迁移至人形机器人控制。我们整理了超过10000小时的以人为中心的原始数据语料库,整合仅视频、仅运动及视频-运动配对数据流,以学习互补的视觉动力学和全身运动学结构。对未来潜在视觉状态和运动的联合预测,促使视觉表示编码未来运动学信息。机器人中期训练将该粗粒度先验适配至机器人视角和身体动力学。在动作后期训练阶段,动作专家通过门控交叉注意力,将来自冻结的适配后先验的视觉预测表示与当前图像及本体感觉相结合,将预测上下文锚定至可执行的全身指令。Being-M0.7在SIMPLE基准测试中取得了对比基线中最高的综合成功率,在真实世界Unitree G1人形机器人移动操作任务上与最强基线表现相当。

英文摘要

Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations. Human video and motion datasets offer scalable supervision, but many contain only video or motion rather than paired video-motion data. Moreover, human motion does not directly specify executable robot actions. We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training. We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure. Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics. Robot mid-training adapts this coarse-grained prior to robot viewpoints and body dynamics. During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands. Being-M0.7 achieves the highest aggregate success rate among the compared baselines on SIMPLE and matches the strongest baseline on real-world Unitree G1 loco-manipulation tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑