发表机构
Institute for AI Industry Research (AIR), Tsinghua University; Institute of Automation, Chinese Academy of Sciences; The Hong Kong University of Science and Technology (Guangzhou); Tsinghua University; School of Information, Renmin University of China; Fudan University; TARS Robotics(清华大学人工智能产业研究院; 中国科学院自动化研究所; 香港科技大学(广州); 清华大学; 中国人民大学信息学院; 复旦大学; TARS Robotics)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TacDyn-WAM提出预测隐式触觉动力学而非重建触觉像素的异构视觉-触觉世界动作模型,通过TacRep目标空间和双专家联合注意力,在UniVTAC和真实机器人任务上达到最先进水平。
AI 中文摘要
世界动作模型通过基于预测的未来状态来调节动作,从而提升机器人操作性能,然而现有的触觉变体大多继承了视频生成流程,通过迭代去噪来重建未来的触觉观测。这种预测在部署漂移下可能变得不可靠:接触位置或力的微小变化可能显著改变触觉像素,即使底层接触演化仍然可预测。我们提出了TacDyn-WAM,一种异构视觉-触觉世界动作模型,它预测隐式触觉动力学,而非重建未来的触觉观测。它学习TacRep,一种动力学感知的触觉目标空间,通过在触觉片段上进行掩蔽时空预测并辅以关系结构蒸馏进行正则化来训练。一个视觉专家和一个隐式触觉动力学专家在各自的目标空间中预测未来的视觉和触觉表征,同时通过联合注意力进行交互;触觉专家在单次前向传播中预测多个时间范围上的未来表征及其变化,并且一个只读触觉记忆提供当前触觉状态。在UniVTAC上,TacDyn-WAM仅使用提供的示范就达到了81.5%的平均成功率,达到了最先进的性能水平,并与在大规模视觉-触觉轨迹上预训练的模型保持竞争力。消融实验证实了触觉通路和TacRep相对于像素重建和静态替代方案的优势。在五个真实机器人任务上,TacDyn-WAM达到了71.0%的平均成功率,而适度规模的预训练将其提升至85.0%,进一步验证了我们的方法。
英文摘要
World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. Such prediction can become unreliable under deployment drift: small changes in contact position or force may substantially alter tactile pixels even when the underlying contact evolution remains predictable. We introduce TacDyn-WAM, a heterogeneous visuo-tactile world action model that predicts implicit tactile dynamics rather than reconstructing future tactile observations. It learns TacRep, a dynamics-aware tactile target space trained through masked spatio-temporal prediction on tactile clips and regularized by relational structure distillation. A visual expert and an Implicit Tactile Dynamics Expert predict future visual and tactile representations in separate target spaces while interacting through joint attention; the tactile expert predicts future representations and their changes at multiple horizons in a single forward pass, and a read-only tactile memory supplies the current tactile state. On UniVTAC, TacDyn-WAM achieves an average success rate of 81.5% using only the provided demonstrations, reaching state-of-the-art-level performance and remaining competitive with models pretrained on large-scale visuo-tactile trajectories. Ablations confirm the benefits of both tactile pathways and TacRep over pixel-reconstruction and static alternatives. On five real-robot tasks, TacDyn-WAM reaches 71.0% average success, and modest-scale pretraining raises it to 85.0%, further validating our method.