发表机构
GensPI; Tsinghua University; BUAA; BIT(GensPI; 清华大学; 北京航空航天大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Motus2,一种用于灵巧操作的自进化通用世界模型,通过模型缩放与数据缩放构建闭环决策学习循环,结合多模态数据与仿生平台实现通用具身智能体的灵巧操作。
AI 中文摘要
通用具身智能体应在统一系统内完成感知、预测、行动、评估与改进。世界模型在构建此类智能体方面展现出巨大潜力,但现有模型通常仅将动作输出头附加到世界模拟器上,未将其耦合为用于策略改进的闭环决策与学习循环。本文提出Motus2,一种用于灵巧操作的自进化通用世界模型。Motus2通过模型缩放与数据缩放推进世界建模:模型缩放方面,一个权重共享的单一模型提供三个控制接口,即策略(世界-动作模型)、模拟器(动作条件世界模型)与评估器(价值模型);策略提出候选动作块,模拟器预测其视觉结果,评估器评估预测结果,三者耦合形成用于策略改进的闭环决策与学习循环,该框架利用精选的专家演示进行动作学习,而失败与次优交互则为动力学建模与价值学习提供有价值的证据。数据缩放方面,Motus2从大规模单目自中心数据推进到同步双目自中心数据,再利用机器人轨迹与补充的人机对齐数据进行机器人领域适配;Motus2进一步研究其滑动窗口上下文的全局自回归与混合记忆扩展,添加用于接触感知控制的触觉反馈,并在具备双目视觉、双臂、双灵巧手与触觉感知的完全仿生平台上实现。总体而言,自中心数据缩放与闭环通用世界模型缩放为实现自进化灵巧操作提供了通用路径。
英文摘要
General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, without coupling them into a closed decision-and-learning loop for policy improvement. We present Motus2, a self-evolving general world model for dexterous manipulation. Motus2 advances world modeling through model scaling and data scaling. For model scaling, a single model with shared weights exposes three control interfaces: a policy (world-action model), a simulator (action-conditioned world model), and an evaluator (value model). The policy proposes candidate action chunks, the simulator predicts their visual consequences, and the evaluator assesses the predicted outcomes. Their coupling forms a closed decision-and-learning loop for policy improvement. This formulation uses curated expert demonstrations for action learning, while failed and suboptimal interactions provide valuable evidence for dynamics modeling and value learning. For data scaling, Motus2 progresses from large-scale monocular egocentric data to synchronized stereo egocentric data, followed by robot-domain adaptation with robot trajectories and supplementary human-robot alignment data. Motus2 further studies global-autoregressive and hybrid-memory extensions of its sliding-window context, adds tactile feedback for contact-aware control, and is instantiated on a fully biomimetic platform with stereo vision, dual arms, dual dexterous hands, and tactile sensing. Together, egocentric data scaling and closed-loop general world model scaling provide a general path toward self-evolving dexterous manipulation.