arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DexTacWAM:用于灵巧操作的视觉-触觉世界-动作模型

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan

arXiv 2609.24976首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; University of California, Berkeley; Northwestern University(伊利诺伊大学厄巴纳-香槟分校; 加利福尼亚大学伯克利分校; 西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DexTacWAM提出视觉-触觉世界-动作模型,通过独立编码指尖并注入触觉潜变量到视频扩散模型,在灵巧操作任务中显著超越基线,并实现数据高效、计算高效的持续触觉适应。

AI 中文摘要

灵巧操作依赖于接触动力学,而这些动力学通常仅从视觉中部分可观测。近期的世界-动作模型(WAMs)将预测性视频世界建模与动作生成相结合,但大多仍以视觉为中心,因此无法直接建模这些接触动力学。我们提出了DexTacWAM,一种视觉-触觉WAM,它独立编码每个指尖,通过手指和姿态感知的触觉压缩器聚合所得特征,并将触觉潜变量注入视频扩散世界模型,以实现联合的视觉-触觉世界建模。在22自由度双臂平台上的六个接触丰富的灵巧操作任务中,DexTacWAM在每个任务上均取得了最高分,平均得分为70.6,而最强基线的平均得分为38.0。消融实验将这一增益归因于将接触演化建模为预测世界状态的一部分,而非仅依赖触觉条件:移除触觉世界建模后,四个任务的平均得分从74.7降至26.6,同时保持相同的触觉特征和动作专家。在冻结的预训练视觉VAE下进行四小时触觉编码器适应后,我们的持续视觉到触觉学习将预训练视频模型扩展到触觉,每个任务仅使用约100个演示,无需触觉中途训练,同时将视觉预测质量保持在仅视觉对应模型的0.5 dB以内。该压缩器保留了融合前接触召回率的89.4%,同时实现了2.26倍更快的训练速度和1.29倍更快的推理速度。这些结果共同表明,预训练视频先验可以以数据和计算高效的方式扩展到分布式多指接触动力学。

英文摘要

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.

Comments22 pages. Project website: https://dextacwam.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑