arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可见触觉:为视觉运动策略渲染接触信息

Visible Touch: Rendering Contact for Visuomotor Policies

Metin Alp Dogan, Edward Sun, Feng Xu, Daniel Wu, Allen Peng, Dennis Hong, Yuchen Cui

arXiv 2609.14156首次发表:更新:

发表机构

University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉运动策略缺乏触觉信息的问题,提出可见触觉方法,将接触信号映射到图像空间,配合开源低成本磁传感器,在LIBERO及真实任务中显著提升策略成功率。

AI 中文摘要

将接触信息整合到视觉运动策略中仍是一个开放问题。触觉对于鲁棒操作至关重要,然而大多数现代策略,包括预训练的视觉-语言-动作(VLA)模型,仅依靠视觉和本体感觉运行。现有弥合这一差距的方法需要专门的触觉硬件、添加独立的触觉编码器,或采用非图像策略主干,这些都与基于预训练2D视觉表示的图像条件策略的现代范式不兼容。我们的关键洞见在于,瓶颈并非接触信息本身,而是其传递方式:当接触信号以策略已关注的场景的同一空间框架呈现时,它们无需架构更改即可被任何图像条件策略直接使用。我们将这一洞见在“可见触觉”中付诸实践,并搭配一个定制的低成本磁接触传感器,该传感器通过参数化CAD到模具的流程,由现成部件制造并开源。在LIBERO基准上,可见触觉在2视图设置中将BC-Transformer的成功率平均提高了15.7个百分点,在1视图设置中也有类似增益;对照比较表明,接触整合策略强烈影响触觉信息的有效利用。这一模式在微调预训练VLA时同样成立:miniVLA在LIBERO上平均提升25个百分点,而π₀.₅在四个真实世界接触密集型任务中,使用我们的定制传感器平均提升30个百分点。

英文摘要

Integrating contact information into visuomotor policies remains an open problem. Touch is essential to robust manipulation, yet most modern policies, including pretrained vision-language-action (VLA) models, operate from vision and proprioception alone. Existing approaches to closing this gap require specialized tactile hardware, add separate tactile encoders, or commit to non-image policy backbones, all incompatible with the modern paradigm of image-conditioned policies built on pretrained 2D visual representations. Our key insight is that the bottleneck is not the contact information itself, but how it is delivered: when contact signals are exposed in the same spatial frame as the scene the policy already attends to, they become directly usable by any image-conditioned policy without architectural changes. We operationalize this insight in Visible Touch, paired with a custom low-cost magnetic contact sensor that is open-sourced and fabricated from off-the-shelf parts via a parametric CAD-to-mold pipeline. Across the LIBERO benchmark, Visible Touch improves BC-Transformer success by 15.7 percentage points on average in the 2-view setting, with similar gains in the 1-view setting; controlled comparisons show that the contact-integration strategy strongly affects how effectively tactile information is used. The pattern holds when fine-tuning pretrained VLAs: miniVLA on LIBERO gains 25 percentage points on average, and $π_{0.5}$ on four real-world contact-rich tasks gains 30 percentage points with our custom sensor.

CommentsAccepted as a Spotlight at the 10th Conference on Robot Learning (CoRL 2026). Project website: https://visibletouch.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑