Dex-X:通过模拟交互从人类视频学习视觉-触觉灵巧操作
Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction
浏览论文内容
中文总结 AI 辅助
DEX-X通过模拟补全触觉信息,从人类视频学习视觉-触觉灵巧操作策略,实现零样本模拟到现实迁移,在抓取和工具使用任务中取得高成功率。
中文摘要 AI 辅助
人类视频是灵巧操作行为的丰富来源,但它们缺乏对于接触丰富的交互至关重要的触觉信息。这引发了一个基本问题:在没有机器人侧数据收集的情况下,机器人能否从人类视频演示中学习可部署的视觉-触觉灵巧操作策略?我们提出了DEX-X,一个通过模拟从人类视频学习视觉-触觉灵巧操作的框架。我们的关键见解是,模拟可以作为触觉补全引擎。给定单目人类演示,DEX-X在模拟中重建手-物体交互,其中物理基础的接触动力学提供了原始视频中不可用的触觉监督。利用这种恢复的触觉信息,我们训练视觉-触觉灵巧操作策略,并将其蒸馏为在点云观测和触觉感知上运行的可部署策略。我们在灵巧手-臂平台上展示了跨多种抓取和接触丰富的工具使用任务的零样本模拟到现实迁移。教师策略在模拟中六个任务类别上实现了65.9%的平均成功率,而蒸馏后的视觉-触觉策略在真实世界立方体拾取上实现了93%的成功率,在具有挑战性的桌面清洁任务上实现了53%的成功率。在物体拾取任务上也观察到了对未见物体几何形状的零样本泛化。我们的结果表明,模拟交互是人类视频与可部署灵巧操作策略之间的关键桥梁,提供了从互联网规模人类视频数据中实现可扩展机器人技能学习所缺失的物理监督。
英文摘要
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection? We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing. We demonstrate zero-shot sim-to-real transfer on a dexterous hand-arm platform across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed on object-picking tasks. Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.
发表机构
- Tsinghua University(清华大学)
- Shanghai Qizhi Institute(上海智己研究院)
- Sharpa(Sharpa公司)
- Tongji University(同济大学)
- Renmin University(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。