arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15060cs.CVcs.RO

EgoTac:从第一视角视觉实现野外环境下的触觉预测

EgoTac: In-the-wild Tactile Prediction from Egocentric Vision

Wenkang Zhang, Chengbo Yuan, Zicheng Zhang, Zhengxue Cheng, Yang Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出EgoTac模型,利用超570万组图像-触觉对训练,可从第一视角视觉预测触觉,在域内及域外测试中性能优于现有方法,能为机器人学习提供触觉先验。

中文摘要 AI 辅助

触觉是灵巧操作的基础,但目前用于机器人学习的大量第一视角人类数据缺乏触觉信息。由于传感器限制,直接收集大规模触觉数据颇具挑战,而人类视频数据丰富、接触信息密集且易于扩展,这催生了一个关键问题:能否仅通过视觉推断出触觉信号?为解决该问题,本文提出EgoTac,这是一种可泛化的模型,能直接从第一视角人类视频中预测丰富的触觉信息。EgoTac在包含超过570万组图像-触觉对的统一数据集上训练,覆盖连续力测量值与二元接触两种类型。通过从该多样化数据集中学习,EgoTac能捕捉各类交互中细微的触觉动态。实验显示其表现优异:域内预测的平均力误差低于0.06N;在域外接触预测基准测试中,EgoTac始终优于当前最优的接触估计器;它还能捕捉真实触觉数据的升降模式,并可对无约束的真实世界视频进行零样本预测。扩展性分析进一步表明,数据的多样性与规模均能稳步提升模型性能。总体而言,EgoTac为从第一视角人类视频中提取触觉先验提供了可扩展的途径,支持广泛适用的感知触觉的机器人学习。

英文摘要

Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce EgoTac, a generalizable model that predicts rich tactile information directly from egocentric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.

发表机构

  • Shanghai Qi Zhi Institute(上海期智研究院)
  • Shanghai Jiao Tong University(上海交通大学)
  • Tsinghua University(清华大学)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑