arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.09565cs.RO

ReTac-ACT:一种基于状态门控的视觉-触觉融合Transformer用于精密装配

ReTac-ACT: A State-Gated Vision-Tactile Fusion Transformer for Precision Assembly

Minchi Ruan, LiangQing Zhou, Hongtong Li, Zongtao Wang, ZhaoMing Lu, Jianwei Zhang, Bin Fang

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出ReTac-ACT,通过双向交叉注意力、体感条件门控网络和触觉重建目标,解决视觉反馈失效下的精密装配问题,实现90%的成功率,优于传统方法。

中文摘要 AI 辅助

精密装配需要在接触密集的'最后毫米'区域进行亚毫米级修正,但视觉反馈因末端执行器和工件遮挡而失效。本文提出ReTac-ACT(Reconstruction-enhanced Tactile ACT),一种视觉-触觉模仿学习策略,通过三种协同机制解决该挑战:(i) 双向交叉注意力实现视觉-触觉特征增强;(ii) 体感条件门控网络在视觉遮挡时动态提升触觉依赖;(iii) 触觉重建目标强制学习与操作相关的接触信息而非通用视觉纹理。在NIST装配任务板M1基准测试中,ReTac-ACT实现90%的圆柱体-孔成功率,显著优于纯视觉和通用基线方法,并在工业级0.1mm间隙下保持80%成功率。消融研究验证了各组件的不可或缺性。ReTac-ACT代码库及包含不同间隙水平的视觉-触觉演示数据集将发布以支持可重复研究。

英文摘要

Precision assembly requires sub-millimeter corrections in contact-rich "last-millimeter" regions where visual feedback fails due to occlusion from the end-effector and workpiece. We present ReTac-ACT (Reconstruction-enhanced Tactile ACT), a vision-tactile imitation learning policy that addresses this challenge through three synergistic mechanisms: (i) bidirectional cross-attention enabling reciprocal visuo-tactile feature enhancement before fusion, (ii) a proprioception-conditioned gating network that dynamically elevates tactile reliance when visual occlusion occurs, and (iii) a tactile reconstruction objective enforcing learning of manipulation-relevant contact information rather than generic visual textures. Evaluated on the standardized NIST Assembly Task Board M1 benchmark, ReTac-ACT achieves 90% peg-in-hole success, substantially outperforming vision-only and generalist baseline methods, and maintains 80% success at industrial-grade 0.1mm clearance. Ablation studies validate that each architectural component is indispensable. The ReTac-ACT codebase and a vision-tactile demonstration dataset covering various clearance levels with both visual and tactile features will be released to support reproducible research.

↑