DexTaG:触觉作为灵巧操作强化学习中的引导
DexTaG: Tactile-as-Guidance in Reinforcement Learning for Dexterous Manipulation
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
- Genesis AI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DexTaG提出触觉引导的强化学习框架,利用手套触觉信号指导灵巧操作策略,训练单一可泛化重定向器并蒸馏为无触觉控制器,在记号笔和锤子任务上优于基线。
AI中文摘要:
基于手套的动作捕捉正成为一种可扩展的灵巧手演示数据收集方法。然而,由于人手与机器人手之间的运动学差异,记录的人体动作无法直接在机器人上执行,尤其是在涉及手内重定向的接触密集型工具使用任务中。先前的工作通过强化学习(RL)或轨迹优化在仿真中弥合了这一差距,但在这种公式下,人类的接触模式难以保留,常常产生不自然的操作和不稳定的功能性抓取。这些方法还为每个参考轨迹训练单独的策略或求解单独的优化,效率低下。为了解决这些问题,我们提出了DexTaG,一个用于灵巧操作的触觉引导强化学习框架。在训练过程中,由手套捕获的触觉信号引导策略搜索朝向测量到的人类接触模式,减少了对精确参考几何形状进行接触监督的依赖。为了提高效率,我们在同一物体的所有训练轨迹上联合训练一个单一的可泛化重定向器。该重定向器进一步蒸馏为一个无触觉的学生控制器,该控制器以目标物体轨迹为条件,用于现实世界部署。在记号笔和锤子操作任务中,DexTaG学习了自然的、接触密集型的行为,而基于距离的接触启发式基线无法学习这些行为,它能够泛化到同一物体和任务的未见过轨迹,并在OakInk2上优于单轨迹基线。
英文摘要:
Glove-based motion capture is emerging as a scalable approach to collecting dexterous-hand demonstration data. However, due to the kinematic gap between the human and robot hand, the recorded human motions cannot be executed directly on the robot, especially for contact-rich tool-use tasks involving in-hand reorientation. Prior work bridges this gap in simulation through reinforcement learning (RL) or trajectory optimization, but the human contact pattern is hard to preserve under such formulations, often producing unnatural manipulation and unstable functional grasps. These methods also train a separate policy or solve a separate optimization for each reference trajectory, which is inefficient. To solve these problems, we propose DexTaG, a tactile-guided RL framework for dexterous manipulation. During training, tactile signals captured by the glove guide policy search toward the measured human contact pattern, reducing reliance on precise reference geometry for contact supervision. To improve efficiency, we train a single generalizable retargeter jointly on all training trajectories of the same object. The retargeter is further distilled into a tactile-free student controller conditioned on the target object trajectory for real-world deployment. On marker-pen and hammer manipulation tasks, DexTaG learns natural, contact-rich behaviors that baselines with distance-based contact heuristics fail to learn, generalizes to held-out trajectories of the same object and task, and outperforms single-trajectory baselines on OakInk2.