发表机构
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发了带均匀RGB照明的视觉-触觉传感器与统一感知框架,采用双头部TimeSformer网络和ResNet-50主干,实现高精度滑移检测与物体分类,为机器人操作提供多模态感知基线。
AI 中文摘要
触觉感知是机器人操作的核心,其中滑移检测是一项典型且关键的任务。然而,现有滑移数据集大多局限于二分类,缺乏细粒度的方向感知。为解决这一局限,我们提出一种配备定制均匀RGB照明的视觉-触觉传感器,以及一套统一感知框架。硬件层面,该传感器实现亚毫米级高精度深度重建;基于此,我们收集了包含15个物体的多任务视觉-触觉数据集,每个样本同步生成深度信息。算法层面,我们设计了双头部TimeSformer网络以处理动态时空滑移。在未见过的物体上,该网络在3类接触状态预测和细粒度8类滑移方向分类上分别达到95.5%和91.5%的鲁棒准确率;此外,基于ResNet-50主干的静态触觉物体分类在15个类别上取得98.8%的出色准确率。所提出的软硬件框架为复杂机器人操作提供高保真反馈与强大的多模态感知基线。
英文摘要
Tactile sensing is central to robotic manipulation, among which slip detection stands out as a quintessential and critical task. However, existing slip datasets are predominantly limited to binary classification, lacking fine-grained directional perception. To address this limitation, we propose a visuo-tactile sensor featuring customized uniform RGB illumination, alongside a unified perception framework. At the hardware level, the sensor achieves high-precision, sub-millimeter depth reconstruction. Based on this capability, we collect a multi-task visuo-tactile dataset encompassing 15 objects, synchronously generating depth information for each data sample. Algorithmically, we design a dual-head TimeSformer network to process dynamic spatiotemporal slip. On unseen objects, this network achieves robust accuracies of 95.5% and 91.5% for 3-class contact state prediction and fine-grained 8-class slip direction classification, respectively. Furthermore, static tactile-based object class recognition utilizing a ResNet-50 backbone yields an outstanding accuracy of 98.8% across 15 categories. The proposed hardware-software framework provides high-fidelity feedback and a powerful multi-modal perception baseline for complex robotic manipulation.
CommentsAccepted at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026. 8 pages, 10 figures