arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAAT:用于数据高效接触丰富型操作的接触感知注意力缩放与触觉掩码方法

CAAT: Contact-Aware Attention Scaling and Tactile Masking for Data-Efficient Contact-Rich Manipulation

Jiaming Jiang, Yuzhe Huang, Hao Liang, Pei Lin, Shengcheng Luo, Fanrong Dong, Jiaping Wu, Chenxi Xiao, Wanlin Li, Ziyuan Jiao

arXiv 2608.01102首次发表:更新:

发表机构

ShanghaiTech University; Beihang University; Beijing Institute for General Artificial Intelligence; Zhejiang University; BUPT(上海科技大学; 北京航空航天大学; 北京通用人工智能研究院; 浙江大学; 北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对接触丰富型操作中视觉-触觉策略缺乏接触感知先验的问题,提出CAAT框架,通过注意力缩放和动态触觉掩码提升策略性能,在仿真和真实实验中均显著优于现有方法。

AI 中文摘要

在接触丰富型操作中,视觉观测主要引导自由空间中的运动,而触觉观测在接触过程中尤为有用。然而,标准的基于Transformer的视觉-触觉策略通常依赖于token拼接或可学习门控,这些方法缺乏显式的接触感知先验,难以从演示中高效学习有效的跨模态表征。为解决这一局限,我们提出CAAT,这是一个轻量型接触感知框架,通过注意力缩放和动态触觉掩码明确融入接触先验。具体而言,CAAT在接触前强调视觉信息,在接触期间强调触觉信息,还通过将当前触觉观测与非接触参考进行比较来抑制静态背景token。CAAT可集成到常用的基于Transformer的策略中,无需修改其动作解码器。在仿真中,将CAAT与ACT集成后,其平均成功率较直接视觉-触觉融合提升18.0个百分点,较门控融合提升10.0个百分点;在使用视觉-触觉UMI平台的真实实验中,CAAT在ACT、Diffusion Policy和π₀上的平均成功率达60.0%,较最强基线平均提升21.1个百分点。这些结果表明,显式接触先验和动态触觉掩码可有效提升不同策略架构的视觉-触觉策略学习及任务性能。

英文摘要

In contact-rich manipulation, visual observations primarily guide motion in free space, whereas tactile observations become particularly informative during contact. However, standard Transformer-based visuo-tactile policies typically rely on either token concatenation or learnable gating. These approaches lack explicit contact-aware priors, making it difficult to efficiently learn effective cross-modal representations from demonstrations. To address this limitation, we propose CAAT, a lightweight contact-aware framework that explicitly incorporates contact priors through attention scaling and dynamic tactile masking. Specifically, CAAT emphasizes visual information before contact and tactile information during contact. It also suppresses static background tokens by comparing the current tactile observation with a non-contact reference. CAAT can be integrated into commonly used Transformer-based policies without modifying their action decoders. In simulation, integrating CAAT with ACT improves the average success rate by 18.0 percentage points over direct visuo-tactile fusion and by 10.0 percentage points over gated fusion. In real-world experiments using a visuo-tactile UMI platform, CAAT achieves an average success rate of 60.0% across ACT, Diffusion Policy, and $π_0$, outperforming the strongest baseline by an average of 21.1 percentage points. These results demonstrate that explicit contact priors and dynamic tactile masking are effective in improving visuo-tactile policy learning and task performance of diverse policy architectures. https://mrjiangjm.github.io/caat/

Comments11 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑