反事实视频生成实现可扩展的人形机器人移动操作
Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
- Amazon FAR(亚马逊FAR)
- UC Berkeley(加州大学伯克利分校)
- Carnegie Mellon University(卡内基梅隆大学)
- Stanford(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出PRISM框架,通过反事实视频生成和接触锚定仿真重建,从少量真实视频扩展训练数据,实现无需真实微调的人形机器人移动操作泛化。
AI中文摘要:
通过视觉模仿教授人形机器人移动操作技能(如搬运多种物体)是通往通用机器人的一条有前景的路径。然而,收集多样化、高质量的交互视频(例如清晰展示人物全身及与物体无遮挡交互的片段)对扩展该方法构成了实际障碍。我们提出PRISM,一个真实-仿真-真实框架,通过将少量真实视频扩充为大规模、多样化的训练集来克服这一限制。PRISM首先通过视频到视频(V2V)生成技术,从少量示例真实视频中生成数百个多样化的“反事实”人-物交互视频。随后,我们的接触锚定真实到仿真流程重建人和物体的运动,将不完美的视频数据重定向为物理上合理的轨迹。这些反事实视频中的类内变异性使我们能够训练一个单一策略,该策略可泛化到每个类别中未见过的物体。我们通过在真实机器人上部署该策略来展示完整流程,且无需任何真实世界微调。仅使用机载深度观测,我们的人形机器人能够拾取、搬运和放下物体,包括箱子、桶、垃圾箱和球,跨越新颖的实例、尺寸和初始配置。
英文摘要:
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.