发表机构
Colorado School of Mines; Northeastern University; Institute for Creative Technologies, University of Southern California; North Carolina State University(科罗拉多矿业大学; 东北大学; 南加州大学创意技术研究所; 北卡罗来纳州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人世界模型对失败不敏感的问题,提出CureWM方法,利用执行验证的反事实失败数据微调模型,显著降低乐观预测并提升成功-失败区分能力。
AI 中文摘要
机器人世界模型支持策略评估、规划和合成数据生成,但这些应用需要能够区分成功动作与失败的预测。在两个架构家族的四个已发布检查点上,我们观察到对动作变化的敏感性较弱,并且对已验证的失败产生类似成功的预测。尽管近期工作将失败纳入模型训练,但哪些数据可以在不改变架构或训练目标的情况下修复已发布的检查点,仍未被充分探索。为此,我们引入CureWM,该方法从成功示范中构建跨越严重性网格的替代动作,通过在仿真或硬件上执行来验证其结果,并在由此产生的失败和幸存的成功以及名义示范上对已发布模型进行微调。这种构建提供了来自共享起始上下文的受控动作对比。在484个保留的LIBERO失败反事实样本上,乐观度从在官方数据上微调后的80%降至四个独立微调的CureWM模型的30%至43%(平均38%)。在物理机器人臂的两次独立评估中,由潜在距离诊断评分为类似成功的失败预测从仅在成功示范上微调后的90%降至使用CureWM的33%。在每任务失败次数、成功重放数据和训练预算匹配的情况下,反事实失败产生的成功-失败价值差距为0.124,而新收集的在线策略失败为0.014。这些发现支持执行验证的反事实重放用于事后修复,并表明为什么减少的乐观度必须与成功-失败区分一起评估。代码可在该https URL获取。
英文摘要
Robot world models support policy evaluation, planning, and synthetic data generation, but these applications require predictions that distinguish successful actions from failures. Across four released checkpoints from two architecture families, we observe weak sensitivity to action changes and success-like predictions on verified failures. Although recent work incorporates failures into model training, which data can repair released checkpoints without changing their architecture or training objective still remains underexplored. To this end, we introduce CureWM, which constructs alternative actions from successful demonstrations across a severity grid, verifies their outcomes through execution in simulation or on hardware, and fine-tunes released models on the resulting failures and surviving successes alongside nominal demonstrations. This construction provides controlled action contrasts from shared starting contexts. On 484 held-out LIBERO failure counterfactuals, optimism falls from 80\% after fine-tuning on the official data to 30--43\% across four independently fine-tuned CureWM models (38\% mean). In two separate evaluations on a physical robot arm, failure predictions scored as success-like by a latent-distance diagnostic decrease from 90\% after fine-tuning on successful demonstrations alone to 33\% with CureWM. With failure counts per task, successful replay data, and training budget matched, counterfactual failures yield a success--failure value gap of 0.124, compared with 0.014 for freshly collected on-policy failures. These findings support execution-verified counterfactual replay for post-hoc repair and show why reduced optimism must be evaluated alongside success--failure discrimination. Code is available at https://github.com/jiuyixu25/CureWM.
Comments23 pages, 4 figures, 12 tables