arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SimpleTouch:视觉-语言-动作模型能否在没有触觉策略预训练的情况下掌握接触丰富的操作?

SimpleTouch: Can Vision-Language-Action Models Master Contact-Rich Manipulation Without Tactile Policy Pretraining?

Chen Yang, Linzhe Shi, Changjie Wu, Hang Zhang, Ronghan Chen, Lingjun Zhang, Xu Hu, Mu Xu, Jiansheng Fan, Chen Wang

arXiv 2610.02784首次发表:更新:

发表机构

Tsinghua University; Amap, Alibaba Group; Hong Kong Polytechnic University(清华大学; 高德,阿里巴巴集团; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SimpleTouch通过向VLA模型添加触觉专家,仅用任务演示进行单阶段训练,无需触觉策略预训练,在UniVTAC和真实任务上均取得最高成功率,证明额外预训练并非必需。

AI 中文摘要

触觉传感为机器人操作提供了关键的接触信息,但将其整合到预训练的视觉-语言-动作(VLA)模型中仍然具有挑战性。一个常见的担忧是,在任务特定的微调期间简单地引入触觉可能无法弥合跨模态差距,导致收益有限甚至成功率下降。因此,现有方法通常依赖大规模触觉策略预训练或单独的视觉-触觉对齐,这增加了数据需求和训练阶段。我们提出了SimpleTouch,一种简单的VLA扩展,通过为$\pi_{0.5}$增加一个触觉专家,来测试这些额外阶段是否必要。利用冻结的预训练触觉编码器的所有令牌,该专家从动作监督和未来触觉潜在状态的多时间跨度预测中学习。这种单阶段训练仅使用任务演示,无需额外的触觉策略预训练或单独对齐。每个任务使用50个演示,SimpleTouch在所有六个UniVTAC任务中取得了评估方法中最高的成功率。其平均成功率达到77.5%,而FTP-$\pi_{0.5}$为45.2%,FTP-1为66.7%,分别对应32.3和10.8个百分点的提升。在四个真实世界任务中,其平均成功率为71.3%,超过FTP-1 8.8个百分点。这些结果表明,在给定预训练的VLA和触觉表示的情况下,额外的触觉策略预训练并非这些任务上取得强性能的先决条件,为接触丰富的操作提供了一条更简单的路径。项目页面:此https URL

英文摘要

Tactile sensing provides essential contact information for robotic manipulation, yet incorporating it into pretrained vision-language-action (VLA) models remains challenging. A common concern is that simply introducing touch during task-specific fine-tuning may fail to bridge the cross-modal gap, yielding limited gains or even reduced success. Consequently, existing methods often rely on large-scale tactile policy pretraining or separate visuotactile alignment, adding data requirements and training stages. We introduce SimpleTouch, a simple VLA extension that augments $π_{0.5}$ with a tactile expert, to test whether these additional stages are necessary. Leveraging all tokens from a frozen pretrained tactile encoder, the expert learns from action supervision and multi-horizon prediction of future tactile latents. This single-stage training uses only task demonstrations, without additional tactile policy pretraining or separate alignment. With 50 demonstrations per task, SimpleTouch achieves the highest success rate among evaluated methods on all six UniVTAC tasks. Its average success rate reaches 77.5%, compared with 45.2% for FTP-$π_{0.5}$ and 66.7% for FTP-1, corresponding to gains of 32.3 and 10.8 percentage points, respectively. Across four real-world tasks, it averages 71.3%, exceeding FTP-1 by 8.8 percentage points. These results demonstrate that, given pretrained VLA and tactile representations, additional tactile policy pretraining is not a prerequisite for strong performance on these tasks, offering a simpler route to contact-rich manipulation. Project page: https://simpletouch-robot.github.io/

Comments27 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑