发表机构
Korea Institute of Science and Technology; Korea Advanced Institute of Science and Technology(韩国科学技术研究院; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出UniDex-ViTac框架,利用人类视频引导模拟生成演示,训练视觉-触觉策略,在模拟和真实机器人上均显著提升灵巧操作成功率。
AI 中文摘要
人类视频提供了灵巧操作的演示,但缺乏机器人可执行的动作和触觉测量。我们提出了UniDex-ViTac,一个利用人类视频引导的模拟来生成机器人演示,并配以指尖接触观测,用于训练可部署的视觉-触觉策略的框架。物体特定的残差强化学习专家将标注的人类-物体交互参考适配到机器人手臂-手系统上。其成功的轨迹将最终机器人动作目标与机器人侧指尖接触观测配对。从十个物体的50个人类演示中,我们收集了10,000条模拟轨迹,用于训练一个基于Transformer的动作分块(ACT)的通用策略。该策略结合了点云、本体感觉和通过指尖标签及单独令牌编码的四个二元接触信号,在部署时无需人类参考或特权物体身份和姿态。接触增强配置在模拟中实现了68.3%的宏平均成功率,而仅点云基线为55.5%。在没有真实机器人演示或策略微调的情况下,它在六个已见和五个未见物体的110次物理试验中成功了73次(66.4%),而基线为60/110(54.5%),提高了11.8个百分点。这些结果支持了从视频引导的模拟交互中学习统一视觉-触觉灵巧操作策略的可行性。项目页面:此https URL
英文摘要
Human videos provide demonstrations of dexterous manipulation but lack robot-executable actions and tactile measurements. We present UniDex-ViTac, a framework that uses human-video-guided simulation to generate robot demonstrations paired with fingertip contact observations for training a deployable visuo-tactile policy. Object-specific residual reinforcement learning specialists adapt annotated human-object interaction references to a robotic arm-hand system. Their successful rollouts pair final robot action targets with robot-side fingertip contact observations. From 50 human demonstrations across ten objects, we collect 10,000 simulated trajectories to train a single Action Chunking with Transformers (ACT) based generalist. The policy combines point clouds, proprioception, and four binary contact signals encoded through fingertip labels and a separate token, without requiring human references or privileged object identity and pose at deployment. The contact-augmented configuration achieves 68.3% macro-average success in simulation, compared with 55.5% for the point-cloud-only baseline. Without real-robot demonstrations or policy fine-tuning, it succeeds in 73/110 physical trials (66.4%) across six seen and five unseen objects, compared with 60/110 (54.5%) for the baseline, an increase of 11.8 percentage points. These results support the feasibility of learning a unified visuo-tactile dexterous manipulation policy from video-guided simulated interactions. Project page: https://unidex-vitac.github.io/
Comments8 pages, 7 figures, 2 tables. Project page: https://unidex-vitac.github.io/