发表机构
Sber AI; Skolkovo Institute of Science and Technology (Skoltech); HSE University; Artificial Intelligence Research Institute (AIRI)(Sber AI; 斯科尔科沃科学技术学院; 高等经济大学; 人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Tactile-JEPA提出一种利用传感器空间拓扑进行自监督预训练的方法,通过双尺度掩码预测学习表示,在力估计和方向估计上显著优于现有方法,提升下游策略学习性能。
AI 中文摘要
触觉感知是机器人执行接触密集、灵巧操作任务时的关键模态,尤其是在视觉遮挡的情况下。虽然预训练的图像编码器在机器人学习流程中是标准配置,但触觉编码器通常仍从原始、含噪的信号从头开始训练,这可能限制了其表达能力。现有的自监督学习(SSL)方法主要聚焦于基于视觉的触觉传感器,而分布式电子皮肤在很大程度上未被涉及。然而,这类传感器具有一个显著特性:其传感元件稀疏且不规则地分布在所覆盖的表面上,这使得直接复用视觉SSL方法并非最优。我们提出了Tactile-JEPA,一种高效的自监督预训练方法,利用触觉传感器的空间排布来学习拓扑感知的表示。具体而言,它被训练为从未被掩码的剩余部分预测被掩码传感元件的嵌入,利用传感器连接图来指导空间掩码。我们的分析表明,有效的触觉表示需要同时捕捉局部接触细节和触觉表面的全局状态,我们通过双尺度掩码实现了这一点。在跨越磁性和压阻式传感器、不同机器人形态以及单传感器和双传感器配置的三个多样化数据集上,Tactile-JEPA相较于先前的最先进方法,将力估计误差降低了6.3%,手内方向误差降低了20.8%,并在其他下游应用(包括策略学习)中持续获得改进。总体而言,我们的结果表明,触觉感知的益处关键取决于编码器预训练的质量,而Tactile-JEPA直接解决了这一问题。代码可在以下网址获取:此https URL。
英文摘要
Tactile sensing is an essential modality for robots performing contact-rich, dexterous manipulation, particularly under visual occlusion. While pre-trained image encoders are standard in robot learning pipelines, tactile encoders are still commonly trained from scratch from raw, noisy signals, which might limit their expressivity. Existing self-supervised learning (SSL) approaches focus predominantly on vision-based tactile sensors, leaving distributed electronic skins largely unaddressed. These sensors, however, have a distinctive property: their sensing elements are sparse and irregularly arranged over the surface they cover, which makes direct reuse of visual SSL methods suboptimal. We present Tactile-JEPA, an efficient self-supervised pre-training method that uses the spatial arrangement of tactile sensors to learn topology-aware representations. Specifically, it is trained to predict the embeddings of masked sensing elements from the unmasked remainder, using the sensor connectivity graph to guide spatial masking. Our analysis shows that effective tactile representations require capturing both local contact details and the global state of the tactile surface, which we achieve through dual-scale masking. Across three diverse datasets spanning magnetic and piezoresistive sensors, different robot embodiments, and single- and paired-sensor configurations, Tactile-JEPA reduces force estimation error by 6.3% and in-hand orientation error by 20.8% over the prior state-of-the-art, with consistent gains in other downstream applications, including policy learning. Overall, our results demonstrate that the benefit of tactile sensing depends critically on the quality of encoder pre-training, a problem which Tactile-JEPA addresses directly. Code is available at https://github.com/E-Kovtun/tactile.