发表机构
Soft Robotic Lab, Department of Mechanical and Process Engineering ETH Zurich, Switzerland(苏黎世联邦理工学院机械与过程工程系软机器人实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Mask2Real-WM,一种用于灵巧操纵的两阶段动作条件世界模型,将像素预测解耦为动力学和渲染模型,利用分割空间较小的模拟到现实差距,通过合成数据预训练和真实演示微调实现各自由度动作可控。
AI 中文摘要
动作条件世界模型可让机器人预测候选动作的未来后果。我们提出Mask2Real-WM,一种用于灵巧操纵的两阶段动作条件世界模型,将像素预测解耦为动力学模型和渲染模型。动力学模型根据过去的掩码和23自由度动作序列预测未来分割掩码,渲染模型使用ControlNet增强的Stable Video Diffusion主干将预测的掩码映射到逼真的RGB。分割空间中较小的模拟到现实差距使动力学模型能够从超过50小时的合成模拟数据上的大规模预训练中受益,然后在少于2.5小时的真实演示上进行微调。在灵巧抓取和放置基准上的实验表明,掩码条件和模拟预训练对于所有23个自由度的每个自由度动作可控性都是必需的。相比之下,整体基线捕获广泛的手部和末端执行器轨迹,但不能可靠地反映细粒度的每个关节动作效果。
英文摘要
Learning action-conditioned world models for dexterous manipulation that are genuinely controllable requires capturing complex, high-dimensional hand kinematics from limited real-world data. We present Mask2Real-WM, a two-stage world model that improves controllability for dexterous hands under a limited real-data budget by decoupling pixel prediction into a dynamics model and a rendering model. The dynamics model predicts future segmentation masks from past masks and a high-dimensional action sequence. The rendering model converts the predicted masks into photorealistic RGB images. This design allows us to train the dynamics model on a large synthetic dataset spanning the full range of hand motions and interactions. We then fine-tune the dynamics model and train the rendering model on only 2.5 h of real demonstrations to obtain a controllable world model. We compare five models with matched training recipes on a 23-DoF robotic system, measuring per-DoF controllability both with random target commands and with sinusoidal per-DoF actuation, judged blind by human raters. The decoupled model achieves the strongest controllability, with further gains from simulation data. Beyond improved controllability, Mask2Real-WM produces sharper frames while maintaining comparable perceptual quality.
Comments8 pages, 7 figures, 3 tables. Preprint. Project page: https://srl-ethz.github.io/Mask2Real-WM/