V2P-Manip:从单目人类视频学习灵巧操作
V2P-Manip: Learning Dexterous Manipulation from Monocular Human Videos
浏览论文内容
中文总结 AI 辅助
提出V2P-Manip框架,从单目人类演示视频中提取具有视觉保真度和物理合理性的轨迹,通过两阶段精炼实现空间对齐与物理一致性,在TACO和OakInk基准上显著优于先前方法。
中文摘要 AI 辅助
实现自主机器人灵巧操作需要大规模精确、类人的动作序列。作为昂贵遥操作数据的可扩展补充,从单目视频中提取兼具视觉保真度和物理合理性的轨迹是具身智能的一个有前景的前沿方向。为此,我们引入V2P-Manip,一个高效的框架,旨在直接从人类演示视频中学习灵巧操作策略。我们建立了一个高效、集成的流水线,涵盖3D资产获取、轨迹估计和灵巧策略学习。为了弥合视觉感知与物理约束之间的差距,我们引入了一个两阶段精炼过程,以强制执行空间对齐和物理一致性。在TACO和OakInk基准上的评估表明,我们的方法在姿态精度、对非结构化环境的适应性以及训练效率方面显著优于先前方法。最终,实验结果证实了在多个合成操作任务上平均成功率超过75%,并验证了提取的操作先验在不同灵巧手形态上的适应性。
英文摘要
Achieving autonomous robotic dexterous manipulation requires precise, human-like action sequences at scale. As a scalable supplement to costly teleoperation data, extracting trajectories with both visual fidelity and physical plausibility from monocular videos represents a promising frontier in embodied AI. To this end, we introduce V2P-Manip, an efficient framework designed to learn dexterous manipulation policies directly from human demonstration videos. We establish an efficient, integrated pipeline encompassing 3D asset acquisition, trajectory estimation, and dexterous policy learning. To bridge the gap between visual perception and physical constraints, we introduce a two-stage refinement process to enforce spatial alignment and physical consistency. Evaluations on the TACO and OakInk benchmarks demonstrate that our approach significantly outperforms previous methods in pose accuracy, adaptability to unstructured environments, and training efficiency. Ultimately, experimental results confirm an average success rate of over 75% across multiple synthetic manipulation tasks and validate the adaptability of the extracted manipulation priors across diverse dexterous hand embodiments.
发表机构
- Zhejiang University(浙江大学)
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。