GE-Act 2.0:面向机器人操作的世界-动作模型的预训练与扩展
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
GE-Act 2.0从零预训练世界-动作模型,结合CoAE、SVP和IDM及KASO优化,在100个任务上无需微调,数据扩展至3万小时使成功率大幅提升,并展现跨具身迁移和指令遵循能力。
中文摘要 AI 辅助
世界-动作模型(WAM)通过预测未来状态来引导机器人动作,从而能够从无动作视频和有动作标签的交互中学习。大多数现有模型继承了预训练的视频生成器,导致WAM的预训练和扩展问题尚未得到充分探索。我们提出了Genie Envisioner Act 2.0(GE-Act 2.0),这是一种世界-动作模型,其可训练生成组件和动作组件均从零开始在操作数据上初始化。它结合了面向控制的自动编码器(CoAE)、单步视觉规划器(SVP)和逆动力学模型(IDM)。CoAE在激进压缩下保留与动作和指令相关的信息,而SVP通过一次可微前向传播生成完整的未来状态,因此视觉规划和逆动力学可以在互补数据上分别预训练。随后,各组件通过知识对齐选择性优化(KASO)进行联合训练,该优化仅选择被判断为与记录动作行为兼容的预测未来,以减少不匹配的监督。我们直接评估预训练检查点,不进行逐任务微调,在20个操作技能组中的100个任务上,使用未见过的场景、背景、光照和物体实例。将协同训练数据从300小时扩展到30,000小时,G1-OP上的成功率从17.1%提升至44.1%,G2-90D上的成功率从13.4%提升至31.1%;尽管G2-90D仅占协同训练数据的不到2%,其成功率提升了17.7个百分点,表明存在跨具身迁移。增益覆盖19/20和18/20的技能组,且技能特定覆盖率与零样本分布外(OOD)成功率强相关(Pearson r=0.80;Spearman rho=0.85)。在相同协议下,模型在至少90%的试验中正确关联物体、颜色、形状和位置引用,并在指令与已承诺行为或常规场景关联冲突时遵循显式指令。
英文摘要
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.
发表机构
- AgiBot Research Team(AgiBot 研究团队)
机构由 AI 辅助整理,请以论文原文为准。