arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.08737cs.RO

Dream-Tac: 用于接触丰富机器人操作任务的统一触觉世界动作模型

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen, Xiaowei Chi, Yaoxu Lyu, Peidong Jia, Sirui Han, Zhihe Lu, Shanghang Zhang

首次发表 更新
浏览论文内容

中文总结 AI 辅助

提出Dream-Tac统一触觉世界动作模型,通过接触门控视觉-触觉融合和接触感知注意力偏置,联合建模动作、未来视觉观察和触觉动态,在六项接触丰富操作任务中平均动作准确率提升31.7%。

中文摘要 AI 辅助

世界动作模型继承了世界模型的预测能力,使得动作生成能够由预期的未来观察引导。然而,它们主要依赖视觉,在接触丰富的操作任务中常常失败,因为关键线索来自物理交互。在本文中,我们提出Dream-Tac,一个统一的触觉世界动作模型,联合建模动作、未来视觉观察和触觉动态。具体来说,Dream-Tac引入了(i)接触门控视觉-触觉融合,以选择性整合触觉信号,以及(ii)接触感知注意力偏置,以更好地调节操作过程中的跨模态交互。为了支持实时部署,我们进一步设计了双级加速策略,在训练期间重新公式化接触感知偏置以保留融合注意力路径,并在推理时引入基于缓存的扩散加速,实现训练速度提升高达2.9倍,推理速度提升1.8倍。在六项接触丰富的操作任务中,Dream-Tac平均动作准确率提升31.7%,证明了统一视觉-触觉世界建模的有效性。代码可在https://github.com/LYFCLOUDFAN/Dream-Tac获取。

英文摘要

World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics. Specifically, Dream-Tac introduces (i) contact-gated visuotactile fusion to selectively integrate tactile signals and (ii) a contact-aware attention bias to better regulate cross-modal interactions during manipulation. To support real-time deployment, we further design a dual-level acceleration strategy, reformulating the contact-aware bias to preserve the fused attention path during training and introducing cache-based diffusion acceleration at inference, achieving up to 2.9$\times$ faster training and 1.8$\times$ faster inference. Across six contact-rich manipulation tasks, Dream-Tac improves action accuracy by 31.7\% on average, demonstrating the effectiveness of unified visuotactile world modeling.Code is available at https://github.com/LYFCLOUDFAN/Dream-Tac.

发表机构

  • Peking University(北京大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Nanjing University(南京大学)
  • State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑