arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.13644cs.ROcs.AIcs.CV

为从人类视频学习灵巧手-物体交互构建世界模型

World Models for Learning Dexterous Hand-Object Interactions from Human Videos

  • FAIR at Meta(Meta 的 FAIR 部门)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

Raktim Gautam Goswami, Amir Bar, David Fan, Tsung-Yen Yang, Gaoyue Zhou, Prashanth Krishnamurthy, Michael Rabbat, Farshad Khorrami, Yann LeCun

更新

AI总结:

本文提出DexWM模型,通过手部关键点提取实现对精细手部动作的建模,提升未来状态预测和零样本迁移能力。

AI中文摘要:

建模灵巧手-物体交互具有挑战性,因为需要理解细微手指运动如何通过与物体接触影响环境。尽管最近的世界模型处理交互建模,但通常依赖于粗粒度的动作空间,无法捕捉精细的灵巧性。因此,我们引入DexWM,一种用于灵巧交互的世界模型,该模型根据过去状态和灵巧动作预测环境的未来潜在状态。为了克服精细标注灵巧数据集稀缺的问题,DexWM使用从第一人称视频中提取的手部关键点表示动作,使模型能够在超过900小时的人类和非灵巧机器人数据上进行训练。进一步,为了准确建模灵巧性,我们发现仅预测视觉特征是不够的;因此,我们纳入了一个辅助手一致性损失,以强制准确的手部配置。DexWM在基于文本、导航或完整身体动作的未来状态预测中优于先前的世界模型,并在Franka Panda臂配以Allegro夹具上展示了对未见技能的强大零样本迁移能力,平均在抓取、放置和触及任务上超过Diffusion Policy超过50%。

英文摘要:

Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically rely on coarse action spaces that fail to capture fine-grained dexterity. We, therefore, introduce DexWM, a Dexterous Interaction World Model that predicts future latent states of the environment conditioned on past states and dexterous actions. To overcome the scarcity of finely annotated dexterous datasets, DexWM represents actions using finger keypoints extracted from egocentric videos, enabling training on over 900 hours of human and non-dexterous robot data. Further, to accurately model dexterity, we find that predicting visual features alone is insufficient; therefore, we incorporate an auxiliary hand consistency loss that enforces accurate hand configurations. DexWM outperforms prior world models conditioned on text, navigation, or full-body actions in future-state prediction and demonstrates strong zero-shot transfer to unseen skills on a Franka Panda arm with an Allegro gripper, surpassing Diffusion Policy by over 50% on average across grasping, placing, and reaching tasks.

↑