发表机构
University of British Columbia; Westlake University; Style3D Research(不列颠哥伦比亚大学; 西湖大学; Style3D研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WorldContact是一种接触中心世界模型,从少量高质量轨迹生成训练数据,以10倍速度模拟购物袋操作,将策略微调后的真实提袋成功率从65%提升至95%。
AI 中文摘要
让机器人适应新物体和新任务需要获取交互经验,而这类经验可能代价高昂。我们提出了WorldContact,一种面向可变形物体操作的接触中心世界模型,它从一组有限的高质量轨迹中构建,以高效地生成额外的训练数据。该模型预测物体动力学时使用的时间步长大于源数值模拟器,后者需要较小的积分步长来解析快速运动并防止相互穿透。我们在16个购物袋操作任务上评估了WorldContact。在单个H100 GPU上的状态展开测量显示,相较于源模拟器,速度提升了10倍(不包括渲染和磁盘I/O)。我们使用生成的数据微调现有的视觉-语言-动作策略,并直接将其部署在真实机器人上。在提袋任务中,仅使用源模拟数据微调的同一策略单次尝试成功率为65%,而使用WorldContact扩展后的数据集微调时,成功率达到了95%。这些结果支持使用WorldContact进行高效数据生成,以促进机器人策略的适应。
英文摘要
Adapting robots to new objects and tasks requires interaction experience that can be costly to obtain. We present WorldContact, a contact-centric world model for deformable-object manipulation, constructed from a limited set of high-quality trajectories to generate additional training data efficiently. It predicts object dynamics using larger time steps than the source numerical simulator, which requires small integration steps to resolve rapid motion and prevent interpenetration. We evaluate WorldContact across 16 shopping-bag manipulation tasks. State-rollout measurements on a single H100 GPU show a $10\times$ speedup over the source simulator, excluding rendering and disk I/O. We use the generated data to fine-tune an existing vision-language-action policy and deploy it directly on a real robot. In bag lifting, the same policy achieves 65% single-attempt success when fine-tuned on source simulation data alone, compared with 95% when fine-tuned on the dataset expanded with WorldContact. These results support efficient data generation with WorldContact for robot policy adaptation.