VisTacAlign:在触觉人类与机器人演示上共同训练灵巧策略
VisTacAlign: Co-Training Dexterous Policies on Tactile Human and Robot Demonstrations
查看机构详情
- ETH Zurich(苏黎世联邦理工学院)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
VisTacAlign通过对齐人类与机器人演示的视觉和触觉模态,共同训练灵巧策略,在精确力任务上优于仅机器人策略。
中文摘要 AI 辅助
人类演示是灵巧操作数据的廉价来源,但将机器人策略与人类演示共同训练需要缩小策略所消费的每种模态中的人类与机器人差距。我们提出了VisTacAlign,一个在人类和机器人演示上共同训练3D视觉-触觉灵巧策略的框架。手套跟踪的人类手部运动通过一次性指尖校正被重定向到17自由度触觉机器人手。然后,人类手部从两个立体视图中被擦除,并用来自机器人记录的像素绘制的姿态化机器人手网格替换,实时立体基础模型在合成图像上重新运行,使得人类点云携带与机器人点云相同的立体误差和可见性。最后,电容式触觉手套在其信号空间中与机器人指尖传感器对齐,提供每个手指可解释的力表示。一个扩散变压器消费点云、本体感觉和每指触觉令牌。在三个需要精确力的真实世界任务——乐高组装、采摘不同大小的草莓以及激活和举起电钻——中,将对齐的人类演示添加到现有机器人数据中优于仅机器人策略,消融实验表明触觉输入和视觉对齐都是必要的。项目页面:此https URL
英文摘要
Human demonstrations are a cheap source of data for dexterous manipulation, but co-training a robot policy on them requires closing the human--robot gap in every modality the policy consumes. We present VisTacAlign, a framework for co-training 3D-visual-tactile dexterous policies on human and robot demonstrations. Glove-tracked human hand motion is retargeted to a 17-DoF tactile robot hand with a one-time fingertip correction. The human hand is then erased from both stereo views and replaced by a posed robot-hand mesh painted with pixels from robot recordings, and a real-time stereo foundation model is re-run on the composite, so the human point clouds carry the same stereo errors and visibility as the robot ones. Finally, a capacitive tactile glove is aligned to the robot's fingertip sensors in its signal space, giving one interpretable per-finger force representation. A diffusion transformer consumes point-cloud, proprioceptive, and per-finger tactile tokens. On three real-world tasks requiring precise force -- Lego assembly, plucking strawberries of varying size, and activating and lifting a power drill -- adding aligned human demonstrations to existing robot data improves over robot-only policies, and ablations show that both tactile input and visual alignment are necessary. Project page: https://vis-tac-align.github.io