通过几何先验预训练提升操作任务的模仿学习效率
Improving Imitation Learning Efficiency for Manipulation through Geometric Prior Pretraining
浏览论文内容
中文总结 AI 辅助
本研究提出利用仅含平面、物体和手的简化几何场景进行预训练,为操作任务模仿学习提供低成本先验,实验表明该方法在模拟和真实机器人上均能提升早期训练成功率。
中文摘要 AI 辅助
将模仿学习策略应用于新的操作任务通常需要收集新的演示数据并重新训练模型,这使得样本效率成为一个实际问题。在大规模机器人数据集上进行预训练在这方面是有效的,但此类数据集的收集和训练成本高昂,而数据增强技术通常需要为每个任务进行新一轮的数据生成和重新训练。一个补充性的问题是:在收集任何任务特定数据之前,能否以极低的成本为策略提供有用的先验。在本研究中,我们构建了一个几何视觉预训练数据集,其中每个场景仅包含一个平面、一个物体和一只手,轨迹自动生成。场景中既没有纹理也没有背景;预训练主要让策略接触手与物体之间的几何关系。此外,将手表示为立方体,避免了对特定机器人形态进行数据集定制。我们使用ACT在三个模拟机器人上分别针对五个操作任务,以及三个真实世界机器人任务上评估了这一几何先验。在许多机器人-任务组合中,与从头训练相比,从几何先验进行微调在训练早期阶段实现了更高的成功率,且仅需使用少量任务演示。这些结果表明,即使是高度简化的几何场景,也能在任务数据有限的情况下,提供一种跨机器人迁移并适用于真实世界任务的有用初始化。
英文摘要
Applying an imitation learning policy to a new manipulation task usually requires collecting new demonstrations and retraining the model, which makes sample efficiency a practical concern. Pretraining on large-scale robot datasets is effective in this respect, but such datasets are costly to collect and train on, while data augmentation techniques typically require a new round of data generation and retraining for each task. A complementary question is what useful prior can be provided to a policy at negligible cost before any task-specific data are collected. In this study, we construct a geometric visual pretraining dataset in which each scene contains only a plane, an object, and a hand, and trajectories are generated automatically. The scenes contain neither textures nor backgrounds; pretraining primarily exposes the policy to the geometric relationship between the hand and the object. Furthermore, representing the hand as a cube avoids tailoring the dataset to a specific robot morphology. We evaluate this geometric prior using ACT on three simulated robots across five manipulation tasks each, as well as on three real-world robot tasks. Across many of these robot--task combinations, fine-tuning from the geometric prior achieves higher success rates in the early stages of training than training from scratch while using only a small number of task demonstrations. These results suggest that even highly simplified geometric scenes can provide a useful initialization that transfers across robots and to real-world tasks when task data are limited.
发表机构
- Waseda University(早稻田大学)
- National Institute of Advanced Industrial Science and Technology (AIST)(产业技术综合研究所(AIST))
- Waseda Research Institute for Science and Engineering(早稻田大学理工学研究所)
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。