发表机构
Shanghai Jiao Tong University; Xense Robotics; BUPT(上海交通大学; Xense机器人公司; 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出VT-MUSE视觉-触觉多模态统一时序表示学习框架,通过两阶段学习解决现有方法的跨模态依赖捕获与时序接触演化忽略问题,在仿真和真实操作任务中性能优于基线
AI 中文摘要
我们提出了VT-MUSE,这是一种面向操作任务的多模态统一时序(Multimodal Unified SEquential)视觉-触觉表示学习框架。现有方法通常在融合前对视觉和触觉观测进行独立编码,这限制了它们捕获细粒度跨模态依赖关系的能力。此外,大多数方法仅关注当前时间步的观测,而忽略了接触的时间演化。VT-MUSE通过两阶段表示学习框架解决了这两个局限性。在第一阶段,通过跨模态时序对齐和掩码视图一致性对模态特定编码器进行联合适配。在第二阶段,条件变分潜在模型将掩码视觉序列与完整触觉历史一同处理。辅助解码器重构被掩码的近期视觉观测并预测触觉深度变化,促使潜在表示同时保留全局视觉上下文和局部接触动态。学习到的表示随后通过门控跨注意力被集成到轻量级Transformer策略中。在仿真基准测试中,VT-MUSE在所有任务上的表现比评估的最强基线高出11个百分点,且在真实世界实验中也取得了显著提升。
英文摘要
We propose VT-MUSE, a Multimodal Unified SEquential representation learning framework for visuotactilemanipulation. Existing approaches often encode visual and tactile observations independently before fusion, limiting their ability to capture fine-grained cross-modal dependencies. Moreover, most methods focus on observations at the current time step and overlook the temporal evolution of contact. VT-MUSE addresses both limitations through a two-stage representation learning framework. In Stage I, modality specific encoders are jointly adapted via cross-modal temporal alignment and masked-view consistency. In Stage II, a conditional variational latent model processes masked visual sequences together with full tactile histories. Auxiliary decoders reconstruct the masked recent visual observations and predict tactile depth changes, encouraging the latent representation to retain both global visual context and local contact dynamics. The learned representation is subsequently integrated into a lightweight Transformer policy through gated cross-attention. On the simulation benchmark, VT-MUSE outperforms the strongest baseline evaluated on all tasks by 11 percentage points and also achieves substantial improvements in real-world experiments.