arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboWorld: 用于通用机器人策略评估的快速可靠神经模拟器

RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation

Byeongguk Jeon, Seonghyeon Ye, JaeHyeok Doo, Sungdong Kim, Minjoon Seo, Hyungmok Son, Kimin Lee

arXiv 2607.01060首次发表:更新:

发表机构

KAIST; Config(韩国科学技术院; Config)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RoboWorld自动化评估流程,结合快速自回归视频世界模型和任务进度感知视觉语言模型评分,通过Step Forcing减少训练-测试不匹配,实现与真实世界评估高度一致。

AI 中文摘要

视频世界模型正成为评估通用机器人策略的可扩展替代方案,绕过了真实世界部署的物理限制和工程负担。然而,使用视频世界模型评估策略仍然具有挑战性,因为世界模型误差可能使生成的轨迹不可靠,且推理速度慢限制了大规模吞吐量。我们引入了RoboWorld,一种自动化评估流程,将快速自回归视频世界模型与任务进度感知的视觉语言模型评分相结合。为了实现可靠的长程自回归世界模型轨迹生成,我们提出了Step Forcing,它结合了锚定和单步自前向上下文,以减少训练-测试不匹配,同时保留动作-观察动态。这些组件共同使RoboWorld能够在不同任务和环境中与真实世界机器人评估高度一致,达到Pearson's r = 0.989和Spearman's ρ = 0.970。

英文摘要

Video world models are emerging as a scalable alternative for evaluating generalist robot policies, bypassing the physical constraints and engineering burdens of real-world deployment. However, evaluating policies with video world models remains challenging, as world-model errors can make generated rollouts unreliable and slow inference limits large-scale throughput. We introduce RoboWorld, an automated evaluation pipeline that pairs a fast autoregressive video world model with a task-progress-aware vision-language model scoring. To enable reliable long-horizon autoregressive world-model rollouts, we propose Step Forcing, which combines anchored and one-step self-forwarded contexts to reduce train-test mismatch while preserving action-observation dynamics. Together, these components enable RoboWorld to align strongly with real-world robot evaluation across tasks and environments, achieving Pearson's r = 0.989 and Spearman's $ρ$ = 0.970.

CommentsProject page: https://byeongguks.github.io/RoboWorld/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑