MAPLE:从第一人称视频中学习的灵巧机器人操作先验知识编码
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
浏览论文内容
中文总结 AI 辅助
MAPLE通过从第一人称视频中学习的先验知识,提升灵巧机器人操作任务的策略学习,有效应对复杂精细的操作需求。
中文摘要 AI 辅助
大规模的第一人称视频数据集捕捉了广泛场景中的人类多样化活动,为理解人类与物体的交互提供了丰富的细节洞察,尤其是那些需要精细灵巧控制的物体。这种复杂的、灵巧的技能需要精确的控制,对于许多机器人操作任务至关重要,但传统数据驱动的方法往往未能充分应对。为解决这一差距,我们利用从大规模第一人称视频数据集中学习到的操作先验知识,以改进灵巧机器人操作任务的策略学习。我们提出了MAPLE,一种新的方法,用于灵巧的机器人操作,该方法学习特征以从第一人称图像中预测物体接触点和接触时刻的详细手部姿态。然后,我们使用这些学习的特征来训练下游操作任务的策略。实验结果表明,MAPLE在4个现有的模拟基准以及一组新设计的4个具有精细物体控制和复杂灵巧技能的挑战性模拟任务中均表现出有效性。MAPLE的优势在使用17自由度灵巧机器人手的现实世界实验中进一步得到强调,而此前的工作在模拟和现实世界实验的同时评估方面一直未被充分探索。我们还展示了模型在第一人称接触点预测任务中的有效性,验证了其在灵巧操作策略学习之外的实用性。
英文摘要
Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grained dexterous control. Such complex, dexterous skills with precise controls are crucial for many robotic manipulation tasks, yet are often insufficiently addressed by traditional data-driven approaches to robotic manipulation. To address this gap, we leverage manipulation priors learned from large-scale egocentric video datasets to improve policy learning for dexterous robotic manipulation tasks. We present MAPLE, a novel method for dexterous robotic manipulation that learns features to predict object contact points and detailed hand poses at the moment of contact from egocentric images. We then use the learned features to train policies for downstream manipulation tasks. Experimental results demonstrate the effectiveness of MAPLE across 4 existing simulation benchmarks, as well as a newly designed set of 4 challenging simulation tasks requiring fine-grained object control and complex dexterous skills. The benefits of MAPLE are further highlighted in real-world experiments using a 17 DoF dexterous robotic hand, whereas the simultaneous evaluation across both simulation and real-world experiments has remained underexplored in prior work. We additionally showcase the efficacy of our model on an egocentric contact point prediction task, validating its usefulness beyond dexterous manipulation policy learning.
发表机构
- ETH Zürich(苏黎世联邦理工学院)
- Mimic Robotics(Mimic机器人公司)
- Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。