arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.11520cs.RO

可扩展的策略评估与视频世界模型

Scalable Policy Evaluation with Video World Models

Wei-Cheng Tseng, Jinwei Gu, Qinsheng Zhang, Hanzi Mao, Ming-Yu Liu, Florian Shkurti, Lin Yen-Chen

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出利用动作条件视频生成模型进行可扩展的策略评估,通过预训练模型利用互联网视频数据,减少现实世界测试需求,提升机器人策略评估效率。

中文摘要 AI 辅助

训练通用性政策用于机器人操作已展现出巨大潜力,因为它们能够在多样化的场景中实现语言条件的多任务行为。然而,评估这些策略仍然困难,因为现实世界测试成本高、耗时且劳动密集。此外,它还需要频繁的环境重置,并且在将未经验证的策略部署到物理机器人上时存在安全风险。手动创建和填充机器人操作的仿真环境资产并未解决这些问题,主要由于所需的大量工程工作和显著的仿真到现实差距,这不仅体现在物理和渲染上。在本文中,我们探索了使用动作条件的视频生成模型作为学习世界模型以进行策略评估的可扩展方法。我们展示了如何将动作条件整合到现有的预训练视频生成模型中。这使得在预训练阶段可以利用互联网规模的野外在线视频,并缓解了需要大量配对视频-动作数据集的需要,而这种数据集对于机器人操作来说是昂贵的。我们的论文考察了数据集多样性、预训练权重和所提出评估流程的常见失败案例的影响。我们的实验表明,在各种指标上,包括策略排名和实际策略值与预测策略值之间的相关性,这些模型为无需现实世界交互的策略评估提供了一种有前途的方法。

英文摘要

Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluating these policies remains difficult because real-world testing is expensive, time-consuming, and labor-intensive. It also requires frequent environment resets and carries safety risks when deploying unproven policies on physical robots. Manually creating and populating simulation environments with assets for robotic manipulation has not addressed these issues, primarily due to the significant engineering effort required and the substantial sim-to-real gap, both in terms of physics and rendering. In this paper, we explore the use of action-conditional video generation models as a scalable way to learn world models for policy evaluation. We demonstrate how to incorporate action conditioning into existing pre-trained video generation models. This allows leveraging internet-scale in-the-wild online videos during the pre-training stage and alleviates the need for a large dataset of paired video-action data, which is expensive to collect for robotic manipulation. Our paper examines the effect of dataset diversity, pre-trained weights, and common failure cases for the proposed evaluation pipeline. Our experiments demonstrate that across various metrics, including policy ranking and the correlation between actual policy values and predicted policy values, these models offer a promising approach for evaluating policies without requiring real-world interactions.

发表机构

  • Nvidia Research(Nvidia 研究院)
  • University of Toronto(多伦多大学)
  • Vector Institute(向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑