arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TACTIC:理解触觉编码器与条件化用于接触丰富机器人操作策略

TACTIC: Understanding Tactile Encoders and Conditioning for Contact-rich Robot Manipulation Policies

Seongjin Bien, Débora Oliveira Makowski, Carlo Kneissl, Reihaneh Mirjalili, Pankhuri Vanjani, Rudolf Lioutikov, Gitta Kutyniok, Florian Walter, Wolfram Burgard

arXiv 2609.30969首次发表:更新:

发表机构

University of Technology Nuremberg; Ludwig-Maximilians University Munich; Karlsruhe Institute of Technology; Deggendorf Institute of Technology; Munich Center for Machine Learning (MCML); University of Tromso; DLR-German Aerospace Center(纽伦堡工业大学; 慕尼黑路德维希-马克西米利安大学; 卡尔斯鲁厄理工学院; 德根多夫理工学院; 慕尼黑机器学习中心; 特罗姆瑟大学; 德国航空航天中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过超过2000次真实世界实验,系统比较了触觉编码器与融合策略,发现不存在普遍最优方案,最佳选择依赖于具体任务。

AI 中文摘要

触觉信息对于机器人中的接触丰富操作任务至关重要。基于视觉的触觉传感器使得设计带有触觉感知的端到端操作策略尤为简便,因为它们能够利用计算机视觉中已有的编码器。然而,这导致了架构、训练数据集和评估协议的极大多样性,使得难以确定哪些设计选择最能编码触觉。在这项工作中,我们弥补了这一空白,并在真实世界实验中针对各种接触丰富操作任务,对触觉编码器和融合策略进行了全面研究。为了实现受控比较,我们在相同的流程和实验设置下训练并评估所有模型,包含超过2000次真实世界 rollout。我们的结果超越了仅比较模拟性能的其他研究,因为模拟性能不一定能转化为真实世界设置,而在真实世界中需要进行大规模评估以获得可靠的统计数据。我们的关键发现是,对于编码视觉-触觉信息,不存在普遍最优的表示或融合策略。相反,最佳的编码器主干和融合方案强烈依赖于具体任务。

英文摘要

Tactile information is essential for contact-rich manipulation tasks in robotics. Vision-based tactile sensors make it particularly easy to design end-to-end manipulation policies with tactile sensing, as they enable the use of existing encoders from computer vision. However, this has led to a huge variety of architectures, training datasets, and evaluation protocols, making it difficult to determine which design choices best encode touch. In this work, we address this gap and present a comprehensive study of tactile encoders and fusion strategies across various contact-rich manipulation tasks in real-world experiments. To enable a controlled comparison, we train and evaluate all models under the same pipeline and experimental setup, comprising more than 2000 real-world rollouts. Our results go beyond other studies that only compare simulation performance, which does not necessarily translate to real-world settings, where large-scale evaluations are needed to obtain reliable statistics. Our key finding is that there is no universally optimal representation or fusion strategy for encoding visual-tactile. Instead, the best encoder backbone and fusion scheme depend strongly on the task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑