arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38172cs.ROcs.CVcs.GR

反事实视频生成实现可扩展的人形机器人移动操作

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

  • Amazon FAR(亚马逊FAR)
  • UC Berkeley(加州大学伯克利分校)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Stanford(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Zihan Wang, Zhen Wu, Pieter Abbeel, Rocky Duan, Jitendra Malik, Carmelo Sferrazza, C. Karen Liu, Guanya Shi, Angjoo Kanazawa

AI总结:

提出PRISM框架,通过反事实视频生成和接触锚定仿真重建,从少量真实视频扩展训练数据,实现无需真实微调的人形机器人移动操作泛化。

AI中文摘要:

通过视觉模仿教授人形机器人移动操作技能(如搬运多种物体)是通往通用机器人的一条有前景的路径。然而,收集多样化、高质量的交互视频(例如清晰展示人物全身及与物体无遮挡交互的片段)对扩展该方法构成了实际障碍。我们提出PRISM,一个真实-仿真-真实框架,通过将少量真实视频扩充为大规模、多样化的训练集来克服这一限制。PRISM首先通过视频到视频(V2V)生成技术,从少量示例真实视频中生成数百个多样化的“反事实”人-物交互视频。随后,我们的接触锚定真实到仿真流程重建人和物体的运动,将不完美的视频数据重定向为物理上合理的轨迹。这些反事实视频中的类内变异性使我们能够训练一个单一策略,该策略可泛化到每个类别中未见过的物体。我们通过在真实机器人上部署该策略来展示完整流程,且无需任何真实世界微调。仅使用机载深度观测,我们的人形机器人能够拾取、搬运和放下物体,包括箱子、桶、垃圾箱和球,跨越新颖的实例、尺寸和初始配置。

英文摘要:

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.

补充信息

相关深度报道

↑