WorldGuide:在潜在世界模型中学习视觉-语言-动作策略的成功-失败边界
WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies
浏览论文内容
中文总结 AI 辅助
WorldGuide通过潜在世界模型学习成功与失败轨迹的区分,结合预测预训练与对比学习,提供可微分奖励优化VLA策略,显著提升可靠性,在LIBERO 100和SimplerEnv上达到最先进性能。
中文摘要 AI 辅助
潜在世界模型通过捕捉动作的后果,为改进视觉-语言-动作(VLA)策略提供了一种有前景的方法。然而,主要基于专家演示训练出的模型对失败结果的接触有限,可能难以区分视觉上相似的成功与失败交互。我们提出WorldGuide框架,在潜在空间中学习这些区别,并利用它们指导策略训练。WorldGuide将成功与失败轨迹上的预测性预训练与匹配成功-失败对的对比学习相结合。学习到的预测器随后提供可微分的奖励,以指导策略和视觉编码器的联合优化。训练后丢弃预测器,因此部署时无需额外的世界模型推理。大量实验表明,WorldGuide显著提高了VLA的可靠性,并在LIBERO 100和SimplerEnv上达到了最先进的性能,分别达到96.8%和72.0%。代码将公开提供。
英文摘要
Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed interactions. We propose \textbf{WorldGuide}, a framework that learns these distinctions in latent space and uses them to guide policy training. WorldGuide combines predictive pretraining on successful and failed trajectories with contrastive learning on matched success--failure pairs. The learned predictor then provides a differentiable reward to guide joint optimization of the policy and visual encoder. The predictor is discarded after training, so deployment requires no additional world-model inference. Extensive experiments show that WorldGuide substantially improves VLA reliability and achieves state of the art performance on LIBERO 100 and SimplerEnv, reaching \textbf{96.8\%} and \textbf{72.0\%}, respectively. Code will be publicly available.
发表机构
- Dalian University of Technology(大连理工大学)
- Beta Infinity
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。