通过双侧力先验实现视觉与触觉感知的对称感知融合用于机器人操作
Symmetry-Aware Fusion of Vision and Tactile Sensing via Bilateral Force Priors for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
本文提出了一种基于双侧力先验的跨模态Transformer,通过结构化注意力机制实现视觉与触觉融合,在机器人操作中实现了96.59%的插入成功率,展示了触觉感知的重要性及物理启发正则化的有效性。
中文摘要 AI 辅助
在机器人操作中,插入任务需要精确且接触丰富的交互,这单靠视觉无法解决。虽然触觉反馈直觉上很有价值,但现有研究显示,简单的视觉-触觉融合往往无法带来一致的改进。在本工作中,我们提出了一种跨模态Transformer(CMT)用于视觉-触觉融合,通过结构化的自注意力和交叉注意力将腕部相机观测与触觉信号结合起来。为了稳定触觉嵌入,我们进一步引入了一种受物理启发的正则化,鼓励双侧力平衡,反映人类运动控制的原则。在TacSL基准测试中,CMT结合对称正则化实现了96.59%的插入成功率,超过了简单和门控融合基线,并接近特权的
英文摘要
Insertion tasks in robotic manipulation demand precise, contact-rich interactions that vision alone cannot resolve. While tactile feedback is intuitively valuable, existing studies have shown that naïve visuo-tactile fusion often fails to deliver consistent improvements. In this work, we propose a Cross-Modal Transformer (CMT) for visuo-tactile fusion that integrates wrist-camera observations with tactile signals through structured self- and cross-attention. To stabilize tactile embeddings, we further introduce a physics-informed regularization that encourages bilateral force balance, reflecting principles of human motor control. Experiments on the TacSL benchmark show that CMT with symmetry regularization achieves a 96.59% insertion success rate, surpassing naïve and gated fusion baselines and closely matching the privileged "wrist + contact force" configuration (96.09%). These results highlight two central insights: (i) tactile sensing is indispensable for precise alignment, and (ii) principled multimodal fusion, further strengthened by physics-informed regularization, unlocks complementary strengths of vision and touch, approaching privileged performance under realistic sensing.
发表机构
- DexAI, Emergent Business Unit, Analog Devices Inc.(DexAI,新兴业务部,安森通公司)
机构由 AI 辅助整理,请以论文原文为准。