arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WM-Craftnet:用于泛化且鲁棒的灵巧手内操作的世界联觉模型

WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation

Jie Yin, Zeyuan Zhao, Xiaojing Tan, Yang Liu, Chiyu Wang, Xinyang Gu

arXiv 2609.07002首次发表:更新:

发表机构

Sharpa Robotics(夏普机器人公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出WM-Craftnet,利用世界联觉模型作为循环任务上下文,结合触觉与深度重建,提升灵巧手内操作的泛化性和鲁棒性,支持多物体及仿真到现实迁移。

AI 中文摘要

泛化且鲁棒的灵巧手内操作要求策略能够从部分且带噪声的观测中推断物体位姿、几何形状、接触状态和潜在滑动。尽管近期基于触觉和视觉-触觉的强化学习方法在受控环境中实现了强手内旋转,但其鲁棒性在姿态偏移、力扰动和物体变化下往往会下降。我们提出WM-Craftnet,一个世界模型条件化框架,从本体感觉、深度、触觉感知和动作中学习紧凑的动作条件潜在动力学,并通过多模态重建和奖励预测进行监督。WM-Craftnet并非将世界模型用于潜在想象或策略优化,而是将学习到的世界联觉模型(WSM)作为非对称Actor-Critic策略的循环任务上下文。重要的是,WSM被训练为从带噪声的深度输入重建干净的深度目标,为真实机器人部署提供去噪的几何状态。对循环基线、辅助头、触觉掩蔽和WSM模态头的消融研究表明,预测性世界建模、干净深度监督和触觉接触线索共同塑造了学习到的状态。在九个沿z轴旋转的物体上预训练的WSM可作为四十九个物体下游策略学习的可复用先验。该上下文改善了多物体旋转,并为未见物体、扰动恢复和仿真到现实迁移提供了定量和定性证据。

英文摘要

Generalizable and robust dexterous in-hand manipulation requires a policy to infer object pose, geometry, contact, and potential slip from partial and noisy observations. Although recent tactile and visuotactile RL methods achieve strong in-hand rotation in controlled settings, their robustness often degrades under pose shifts, force disturbances, and object variation. We propose WM-Craftnet, a world-model-conditioned framework that learns compact action-conditioned latent dynamics from proprioception, depth, tactile sensing, and actions, supervised by multimodal reconstruction and reward prediction. Rather than using the world model for latent imagination or policy optimization, WM-Craftnet uses the learned World Synesthesia Model (WSM) as recurrent task context for an asymmetric actor--critic policy. Importantly, WSM is trained to reconstruct clean depth targets from noisy depth inputs, providing a denoised geometric state for real-robot deployment. Ablations over recurrent baselines, auxiliary heads, tactile masking, and WSM modality heads show that predictive world modeling, clean-depth supervision, and tactile contact cues all shape the learned state. A WSM pretrained on nine \(z\)-axis objects serves as a reusable prior for \(49\)-object downstream policy learning. This context improves multi-object rotation, with quantitative and qualitative evidence for unseen-object, perturbation-recovery, and sim-to-real transfer.

CommentsAccepted to CoRL2026. Project website: https://wmcraftnet.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑