发表机构
HKUST (GZ); Xspark AI; PKU; THU; HKU(香港科技大学(广州); Xspark AI; 北京大学; 清华大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DexTouch-WM,利用人类触觉数据训练动作条件世界模型,预测灵巧操作的视觉与触觉动态,通过人类交互扩展显著提升机器人预测性能。
AI 中文摘要
学习接触丰富的灵巧操作预测模型需要密集的触觉交互,但此类数据在真实机器人上扩展成本高昂,且受限于特定于具体传感器的设备。我们提出了DexTouch-WM,一种动作条件世界模型,它从可扩展的人类触觉中学习,以联合预测未来的RGB观测和双侧触觉动力学。我们的洞察在于,当人类和机器人的触觉观测及动作空间变得兼容时,它们共享可迁移的接触动力学。我们在人类和灵巧机器人手上部署了具有共享传感布局的柔性压阻阵列,并将人类动作重定向到机器人动作空间,从而使人类交互能够监督用于真实机器人预测的同一动力学模型。DexTouch-WM通过解剖感知触觉标记和对齐的动作条件,将预训练的视频专家与轻量级触觉专家耦合。在从人类到机器人的扩展实验中,我们保持五小时的真实机器人监督固定,同时将人类交互从0小时增加到100小时,并观察到在保留的机器人域视觉、几何和接触预测方面有显著改进,尽管人类和机器人任务集不重叠。在预测之外,我们将世界模型评估为用于策略评估的替代环境,以及用于真实机器人策略学习的合成轨迹生成器,表明可扩展的人类交互为学习灵巧机器人世界模型提供了互补的数据轴。
英文摘要
Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.
CommentsAccept to IROS 2026 Workshop RoBoWoMo (Lightning Talk)