AnyWorld:用于跨 embodiments 泛化的分解式第一人称世界模型
AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization
浏览论文内容
中文总结 AI 辅助
AnyWorld 框架通过分解交互为动作、相机、embodiment 因子,将人类交互重组为机器人原生经验,生成数据可提升机器人操作性能,验证了动作校准与视觉重组的必要性。
中文摘要 AI 辅助
大规模采集接触丰富的机器人经验仍是实现可泛化操作的主要瓶颈。除数据量外,机器人学习还需要跨 embodiments、视角和场景的多样化经验。人类第一人称视频提供了丰富的物理交互,但每个视频仅捕获单一身体、相机轨迹和环境下的窄范围经验。我们提出 AnyWorld,一种跨 embodiment 的世界建模框架,无需配对的人机演示即可将单个人类交互扩展为多样化的机器人原生 rollout。我们的模型将交互分解为动作、相机和 embodiment:动作控制捕获运动结构,相机控制指定视角演化,目标 embodiment 上下文定义作用身体及其交互几何。该公式允许独立重组 embodiment、视角和场景因子,使单个模型能生成大量机器人域经验,同时保留底层动力学和对象交互。我们通过大规模人类交互预训练,随后进行混合 embodiment 微调来训练模型。实验表明,我们的模型支持跨 embodiments、视角和场景的可控重组,我们进一步证明生成的数据可提升 RoboCasa GR1 桌面基准和真实 IRON 人形机器人的操作性能。除整体提升外,我们还测试了未配对人类经验是否可重组为针对策略差距的机器人原生视频-动作对:可控 IRON 干预修正了虚假完成先验并建立了语言接地的空间目标选择;仅动作的反事实干预无法可靠学习后者,表明动作校准和视觉重组均为必要。
英文摘要
Collecting contact-rich robot experiences at scale remains a major bottleneck for generalizable manipulation. Beyond data quantity, robot learning also requires diverse experiences across embodiments, viewpoints, and scenes. Human egocentric videos provide abundant physical interactions, but each video captures only a narrow slice of experience under a single body, camera trajectory, and environment. We propose AnyWorld, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations. Our model factorizes an interaction into action, camera, and embodiment: action controls capture the motion structure, camera controls specify viewpoint evolution, and the target embodiment context defines the acting body and its interaction geometry. This formulation enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain experiences while preserving the underlying dynamics and object interactions. We train the model with large-scale human interaction pretraining followed by mixed-embodiment fine-tuning. Experiments show that our model supports controllable recomposition across embodiments, viewpoints, and scenes, and we further demonstrate that the generated data can improve manipulation performance on the RoboCasa GR1 tabletop benchmark and a real IRON humanoid robot. Beyond aggregate gains, we test whether unpaired human experience can be recomposed into robot-native video-action pairs that target a policy gap. Controlled IRON interventions correct a spurious completion prior and establish language-grounded spatial target selection; an action-only counterfactual intervention fails to learn the latter reliably, showing that both action calibration and visual recomposition are necessary.
发表机构
- Nanyang Technological University(南洋理工大学)
- Institute of Advanced Intelligence and Computing, A*STAR(新加坡科技研究局高级智能与计算研究院)
- XPENG Robotics(小鹏机器人)
- Zhejiang University(浙江大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。