arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

灵巧触觉世界模型

Dexterous Tactile World Model

Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan, Zhengxiang Yu, Fengyu Yang, Tianyu Liu, Zhiwen Fan, Daniel Rakita

arXiv 2609.34286首次发表:更新:

发表机构

Yale University; Texas A&M University; University of California, Los Angeles; University of Washington(耶鲁大学; 德克萨斯A&M大学; 加州大学洛杉矶分校; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出灵巧触觉世界模型(DTWM),融合视觉与触觉信号预测操作未来帧,减少手部运动低估并提升预测精度。

AI 中文摘要

用于操作的世界模型通常从视频中训练,然而决定操作如何展开的事件,例如建立接触和释放接触,难以通过视觉观察,且通常更容易通过触觉感知。我们提出了灵巧触觉世界模型(DTWM),一种视频世界模型,用于从观察到的视频和每只手上佩戴的手套所获取的触觉信号,对自我中心操作进行未来帧预测。我们通过在手部视频令牌对应位置处的零初始化残差,将预训练的视频扩散变换器条件化于每只手的触觉信号,同时因果掩码防止预测帧访问未来信息。与在架构、参数和训练上匹配的仅视觉模型相比,DTWM 将手部运动的低估从 23% 降低到 9%,同时在每个模型的三次训练运行中,手部区域的感知误差降低了 7.4%。这种优势也随预测时域的增加而增大,在较后的预测块中的改进比第一个块大约大 4.1 倍。DTWM 在相同设置下也优于其他视觉-触觉世界模型,并且使用触觉进行训练即使在推理时没有触觉也能改善未来帧预测。消融研究表明,模型受益于力的大小和空间位置:用二元接触状态(无论是每只手还是每个位置)替换触觉信号会增加预测误差。力的观察过程指示交互将持续还是改变。

英文摘要

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.

CommentsProject page: https://adonis-galaxy.github.io/dtwm-project-page/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑