arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05588cs.ROcs.CV

GE-Act 2.0:面向机器人操作的世界-动作模型的预训练与扩展

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

AgiBot Research Team, Renhang Liu, Wenzhi Zhao, Zhuo Yang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Guanghui Ren, Youlun Peng, Rongjun Jin, Nan Wang, Sukai W… 展开作者

AgiBot Research Team, Renhang Liu, Wenzhi Zhao, Zhuo Yang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Guanghui Ren, Youlun Peng, Rongjun Jin, Nan Wang, Sukai Wang, Xindong He, Jinyuan Feng, Ziyu Xiong, Linqing Zhong, Yifei Wei, Feng Han, Long Zhang, Da Huang, Nanshu Zhao, Chenghao Yin, Mo Wu, Zhaodong Yan, Kongtao Hu, Yuxiang Yan, Aogelijiang Niyazi, Yu Fang, Jia Zeng, Lizhu Meng, Daizhen Lv, Haoyu Cao, Zhiwen Hou, Lianjin Ye, Yuehan Niu, Zhikai Cai, Xuan Hu, Hui Min, Xiongfeng Cai, Yue Liao, Jing Wu, Soujanya Poria, Ye Li, Sanping Zhou, Maoqing Yao

首次发表
浏览论文内容

中文总结 AI 辅助

GE-Act 2.0从零预训练世界-动作模型,结合CoAE、SVP和IDM及KASO优化,在100个任务上无需微调,数据扩展至3万小时使成功率大幅提升,并展现跨具身迁移和指令遵循能力。

中文摘要 AI 辅助

世界-动作模型(WAM)通过预测未来状态来引导机器人动作,从而能够从无动作视频和有动作标签的交互中学习。大多数现有模型继承了预训练的视频生成器,导致WAM的预训练和扩展问题尚未得到充分探索。我们提出了Genie Envisioner Act 2.0(GE-Act 2.0),这是一种世界-动作模型,其可训练生成组件和动作组件均从零开始在操作数据上初始化。它结合了面向控制的自动编码器(CoAE)、单步视觉规划器(SVP)和逆动力学模型(IDM)。CoAE在激进压缩下保留与动作和指令相关的信息,而SVP通过一次可微前向传播生成完整的未来状态,因此视觉规划和逆动力学可以在互补数据上分别预训练。随后,各组件通过知识对齐选择性优化(KASO)进行联合训练,该优化仅选择被判断为与记录动作行为兼容的预测未来,以减少不匹配的监督。我们直接评估预训练检查点,不进行逐任务微调,在20个操作技能组中的100个任务上,使用未见过的场景、背景、光照和物体实例。将协同训练数据从300小时扩展到30,000小时,G1-OP上的成功率从17.1%提升至44.1%,G2-90D上的成功率从13.4%提升至31.1%;尽管G2-90D仅占协同训练数据的不到2%,其成功率提升了17.7个百分点,表明存在跨具身迁移。增益覆盖19/20和18/20的技能组,且技能特定覆盖率与零样本分布外(OOD)成功率强相关(Pearson r=0.80;Spearman rho=0.85)。在相同协议下,模型在至少90%的试验中正确关联物体、颜色、形状和位置引用,并在指令与已承诺行为或常规场景关联冲突时遵循显式指令。

英文摘要

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.

发表机构

  • AgiBot Research Team(AgiBot 研究团队)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑