arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34182cs.ROcs.AI

基于人类示教的统一视觉-触觉-动作建模用于灵巧操作

Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation

Wenqiao Li, Qianyou Zhao, Jiawen Hao, Xuezhou Zhu, Tengyu Liu, Kaifeng Zhang, Chuan Wen, Siyuan Huang

首次发表
浏览论文内容

中文总结 AI 辅助

提出统一视觉-触觉-动作模型,利用人类触觉示教数据提升灵巧操作策略,在五个任务上实现70%平均成功率,显著优于基线。

中文摘要 AI 辅助

灵巧操作需要触觉反馈。然而,机器人触觉示教难以规模化,因为灵巧手遥操作向操作员提供的触觉反馈有限。相比之下,人类示教提供了更多样化触觉交互的可扩展来源。我们基于一个简单的前提:手可以改变,但交互的基本物理规律不变。我们利用人类触觉数据来改进灵巧操作策略。具体而言,我们首先构建了一个触觉动作捕捉系统,同步记录图像、触觉信号和手部运动。利用该系统,我们构建了UVTA数据集,涵盖五个接触密集任务,包含1,000个人类示教,覆盖多样化的交互模式,以及每个任务150个机器人示教。为了将人类交互的基本物理规律迁移到机器人控制,我们提出了统一视觉-触觉-动作模型,将两种具身映射到对齐的触觉和动作表示,并联合预测未来的动作和触觉轨迹。联合目标使人类示教能够监督接触感知表示学习,而在部署期间仅执行机器人动作。在五个任务的实际机器人评估中,我们的方法实现了70%的平均成功率,优于最强的视觉-触觉基线(29%)和架构消融(42%)。性能随额外人类示教持续提升,且在每任务1,000个示教时未出现饱和,验证了可扩展人类触觉数据对灵巧操作的有效性。项目页面可在此URL获取。

英文摘要

Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • State Key Laboratory of General Artificial Intelligence, BIGAI(北京通用人工智能研究院 通用人工智能全国重点实验室)
  • Sharpa Robotics
  • Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑