世界合一演示:用于学习开放世界移动操作的合成数据引擎
Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation
AI总结:
研究如何通过合成数据引擎WANDA从单个演示学习开放世界移动操作策略,利用重建背景、重排交互片段、校正状态扩展等方法,实现长期鲁棒性、空间泛化和跨环境泛化,还支持跨实体数据生成。
AI中文摘要:
学习开放世界移动操作策略需要大量数据以实现空间泛化、长期鲁棒性和场景泛化。当前流行的数据收集范式,如遥操作和UMI,在大规模应用时需要高昂的人力和成本。为突破手动数据收集的限制,我们试图通过可扩展的数据生成来最大化每个人类演示的价值。为此,我们引入了WANDA:通过合成数据引擎从单个演示中学习开放世界移动操作。WANDA首先从源RGBD观测中重建背景高斯点云和机器人与物体的交互轨迹,作为后续规划和渲染的世界基础。然后,它将富含接触的机器人与物体交互片段重新排列成广泛的空间配置,利用全身运动规划将它们链接成新的轨迹。为增强长期鲁棒性,它应用校正状态扩展来增加移动操作不同阶段的机器人和物体状态多样性。为实现跨环境泛化,轨迹在从日常照片生成的各种3D世界上合成。此外,我们通过将渲染的机器人和物体网格与高斯点云背景合成来生成逼真的观测。我们在各种场景的广泛模拟和现实世界任务中评估了我们的方法。实验表明,用WANDA训练的策略从一个真实演示中实现了长期鲁棒性、广泛的空间泛化和跨环境泛化。此外,WANDA自然支持跨实体数据生成,并在另一个形态不同的移动操纵器上进行零样本部署验证。
英文摘要:
Learning open-world mobile manipulation policies requires vast data to achieve spatial generalization, long-horizon robustness, and scene generalization. Current prevailing data collection paradigms, teleoperation and UMI, demand prohibitive human effort and cost at scale. To scale beyond the limits of manual data collection, we seek to maximize the value of each human demonstration by scalable data generation. To this end, we introduce WANDA: learning open-World mobile mANipulation from one demonstration via a synthetic DAta engine. WANDA first reconstructs background Gaussian splats and robot-object interaction trajectories from source RGBD observations, as a world substrate for later planning and rendering. It then rearranges contact-rich robot-object interaction segments into extensive spatial configurations, utilizing whole-body motion planning to chain them into new trajectories. To enhance long-horizon robustness, it applies Corrective State Expansion to increase the robot and object state diversity at different stages of mobile manipulation. To unlock cross-environment generalization, trajectories are synthesized on diverse generated 3D worlds from everyday photos. Furthermore, we synthesize photo-realistic observations by compositing rendered robot and object meshes with Gaussian splatting backgrounds. We evaluate our approach on extensive simulation and real-world tasks in various scenes. Experiments show that policies trained with WANDA achieve long-horizon robustness, broad spatial generalization and cross-environment generalization from one real demonstration. Moreover, WANDA naturally supports cross-embodiment data generation, validated by zero-shot deployment on another mobile manipulator with a distinct morphology.