arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02503cs.RO

VT-WAM: 面向密集接触操作的视觉-触觉世界动作模型

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

Shuai Tian, Yupeng Zheng, Yuhang Zheng, Songen Gu, Yujie Zang, Yuxing Qin, Weize Li, Haoran Li, Wenchao Ding, Dongbin Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

提出VT-WAM,在统一流匹配框架中联合学习视觉预测、触觉变形预测和动作预测,通过非对称混合Transformer注意力和接触门控注意力引导,在六项真实密集接触操作任务中平均成功率71.67%。

中文摘要 AI 辅助

密集接触操作需要策略对局部变形、压力、滑动和摩擦做出反应,但这些线索在时间上稀疏且在视觉观察中通常不可见。现有的视觉-触觉策略通常直接将触觉观测输入动作预测,但很少在动作生成过程中建模触觉变形动态。本文提出VT-WAM,一种视觉-触觉世界动作模型,在统一的流匹配框架中联合学习未来视觉预测、触觉变形预测和动作预测。具体地,VT-WAM引入了(1)非对称混合Transformer注意力,以桥接首帧视觉锚点与时间触觉动态,以及(2)接触门控动作-视觉-触觉注意力引导,以鼓励动作查询在接触阶段依赖触觉证据。在六项真实世界密集接触操作任务中,VT-WAM实现了71.67%的平均成功率,比Fast-WAM高出26.67%,比OmniVTLA高出35.84%。消融实验表明,建模触觉变形动态和引导接触阶段触觉注意力对于密集接触任务都很重要。项目网站:此https URL。

英文摘要

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observations. Existing visual-tactile policies usually feed tactile observations directly into action prediction, but rarely model tactile deformation dynamics during action generation. In this paper, we introduce VT-WAM, a Visual-Tactile World Action Model that jointly learns future visual prediction, tactile deformation prediction, and action prediction within a unified flow matching framework. In particular, VT-WAM introduces (1) Asymmetric Mixture-of-Transformers (MoT) attention to bridge a first-frame visual anchor with temporal tactile dynamics, and (2) contact-gated Action-Visual-Tactile Attention Guidance (AVTAG) to encourage action queries to rely on tactile evidence during contact phases. Across six real-world contact-rich manipulation tasks, VT-WAM achieves a 71.67% average success rate, outperforming Fast-WAM by 26.67% and OmniVTLA by 35.84%. Ablations demonstrate that modeling tactile deformation dynamics and guiding contact-phase tactile attention are both important for contact-rich tasks. Project website: https://vt-wam.github.io/.

发表机构

  • SKL-MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统管理与控制国家重点实验室)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • TARS Robotics
  • National University of Singapore(新加坡国立大学)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑