AI 中文总结
本文提出Grasp2Twist,一种基于强化学习的双手灵巧开瓶系统,通过几何包围度量、统一策略和三阶段课程解决多阶段任务、持续旋转及仿真到现实迁移,实现零样本迁移并达到88%成功率。
AI 中文摘要
本文提出了Grasp2Twist,一个双手灵巧操作系统,通过强化学习学习抓取并旋转打开瓶盖。学习这一任务面临三个挑战:为多阶段任务学习统一策略、维持瓶盖旋转以及仿真到现实的迁移。为应对第一个挑战,我们引入了一个连续包围度量来引导抓取形成,以及一个二值包围指示器来引导从抓取到旋转的过渡,以实现统一策略学习。两者均基于物体中心与手部手掌和指尖形成的凸包之间的几何关系推导而来。运动学约束限制了手在固定接触下旋转瓶盖的角度,因此持续旋转需要手指接触的重新配置。我们采用三阶段课程来促进这些接触变化的探索,并增强仿真到现实迁移的鲁棒性。通过我们的方法,学习到的策略展示了手指步态,重新配置手-物接触以维持瓶盖旋转。该策略零样本迁移到物理系统,并在六个家用容器(包括花生酱、维生素和速溶咖啡罐)上实现了88%的任务成功率。消融实验进一步验证了几何包围在抓取形成中的作用以及课程在接触重新配置探索中的作用。
英文摘要
This paper presents Grasp2Twist, a bimanual dexterous manipulation system that learns to grasp and twist open jar lids using reinforcement learning. Learning this task raises three challenges: learning a unified policy for a multi-stage task, sustaining lid twisting, and sim-to-real transfer. To address the first challenge, we introduce a continuous enclosure measure to guide grasp formation and a binary enclosure indicator to guide the grasp-to-twist transition for unified policy learning. We derive both from the geometric relationship between the object center and the convex hull formed by the hand's palm and fingertips. Kinematic constraints limit how far the hand can rotate the lid with fixed contacts, so sustained twisting requires finger contact reconfiguration. We use a three-stage curriculum to facilitate exploration of these contact changes and also improve robustness for sim-to-real transfer. With our approach, the learned policy demonstrates finger gaiting, reconfiguring hand-object contacts to sustain lid rotation. It transfers zero-shot to the physical system and achieves an 88% task success rate across six household containers, including peanut-butter, vitamin, and instant-coffee jars. Ablations further validate the roles of the geometric enclosure in grasp formation and the curriculum in contact-reconfiguration exploration.
Comments9 pages, 6 figures. Under review