arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10288cs.CV

TouchScale:用于视觉-触觉学习的500小时人类视觉与触觉数据

TouchScale: 500 Hours of Human Vision and Touch for Visual-Tactile Learning

Dayou Li, Hao Wang, Qianqian Yang, Zihao Zhu, Haoquan Fang, Ziyao Zeng, Yan Han, Zihan Wang, Yan Wang, Baoru Huang, Dilin Wang, Kenji Shimada, Yiyue Luo, Manlin… 展开作者

Dayou Li, Hao Wang, Qianqian Yang, Zihao Zhu, Haoquan Fang, Ziyao Zeng, Yan Han, Zihan Wang, Yan Wang, Baoru Huang, Dilin Wang, Kenji Shimada, Yiyue Luo, Manling Li, Teresa Lv, Mustafa Mukadam, Rakesh Ranjan, Ruohan Zhang, Qi He, Changliu Liu, Xu Chen, Marco Pavone, Bangya Liu, Jiachen Li, Masayoshi Tomizuka, Zhiwen Fan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出TouchScale,一个500小时统一传感器采集的人类视觉-触觉数据集,显著提升零样本触觉预测和机器人操作成功率,证明大规模一致传感数据对感知和操作的价值。

中文摘要 AI 辅助

大规模以自我为中心的人类交互数据正成为具身学习物理监督的重要来源,然而仅视频无法记录表征物理交互的接触与压力。近期的视觉-触觉数据集提供了这一缺失的监督信息,但其同步触觉数据的规模远小于人类视频。此外,最大的数据资源往往合并来自不同传感器或标注流程的记录,这使得数据规模的影响难以单独评估。因此,我们引入了TouchScale,一个使用单一统一可穿戴设备记录的500小时接触丰富的人类交互数据集。其约2000个预定义任务描述涵盖日常活动和结构化操作,每条记录在时间上对齐了以自我为中心的RGB-D视频、腕部RGB视频以及密集的全手双手触觉测量。与先前的触觉数据相比,在完整TouchScale上训练将未见触觉传感器数据的零样本接触IoU从0.134提升至0.383。在TouchScale上预训练视觉编码器也在三个基准上取得了最高的动作识别准确率,优于所比较的视觉-触觉数据集。用于机器人策略的视觉-触觉中期训练时,TouchScale将四个接触丰富操作任务的平均真实世界成功率从22.5%提升至57.5%。在传感器和采集协议固定的情况下,随着使用更多TouchScale数据,零样本触觉预测和机器人成功率均呈现整体上升趋势。这些结果表明,以一致传感方式大规模收集的人类视觉-触觉数据有益于感知和机器人操作。我们将公开发布TouchScale,包括所有同步的视觉-触觉记录和重建的物体模型,以支持未来可扩展视觉-触觉学习的研究。

英文摘要

Large-scale egocentric human interaction data is becoming an important source of physical supervision for embodied learning, yet video alone leaves the contact and pressure that characterize physical interaction unrecorded. Recent visual-tactile datasets provide this missing supervision, but their synchronized tactile data remain far smaller in volume than human video. Moreover, the largest resources often merge recordings from different sensors or annotation procedures, which makes the effect of data scale difficult to isolate. We therefore introduce TouchScale, a 500-hour dataset of contact-rich human interaction recorded with a single unified wearable setup. Its approximately 2K predefined task descriptions span everyday activities and structured manipulation, and each recording temporally aligns egocentric RGB-D video with wrist RGB video and dense full-hand bimanual tactile measurements. Compared with prior tactile data, training on the full TouchScale raises zero-shot contact IoU on data from an unseen tactile sensor from 0.134 to 0.383. Pretraining a visual encoder on TouchScale also yields the highest action recognition accuracy on three benchmarks among the compared visual-tactile datasets. Used for visual-tactile mid-training of a robot policy, TouchScale improves the average real-world success rate across four contact-rich manipulation tasks from 22.5% to 57.5%. With the sensor and collection protocol held fixed, both zero-shot tactile prediction and robot success show an overall upward trend as more TouchScale data is used. These results suggest that human visual-tactile data collected at scale with consistent sensing benefits both perception and robot manipulation. We will publicly release TouchScale, including all synchronized visual-tactile recordings and reconstructed object models, to support future research on scalable visual-tactile learning.

发表机构

  • Texas A&M University(德克萨斯A&M大学)
  • Google DeepMind(谷歌DeepMind)
  • CMU(卡内基梅隆大学)
  • Stanford University(斯坦福大学)
  • Yale University(耶鲁大学)
  • Microsoft(微软)
  • Overfit Lab(Overfit实验室)
  • NVIDIA(英伟达)
  • University of Liverpool(利物浦大学)
  • Meta
  • University of Washington(华盛顿大学)
  • Northwestern University(西北大学)
  • Sony(索尼)
  • Georgia Tech(佐治亚理工学院)
  • UC Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑