发表机构
Cardinal AI Lab; University of California, Berkeley; Hong Kong University of Science and Technology; National University of Singapore; HPC-AI Lab(卡迪纳尔人工智能实验室; 加州大学伯克利分校; 香港科技大学; 新加坡国立大学; 高性能计算与人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文指出扩展世界模型的传统策略效率低,提出将游戏开发作为可验证轨迹数据引擎,构建RLHEV后训练范式,为空间世界模型提供高质量奖励信号与长程轨迹数据。
AI 中文摘要
扩展世界模型的常用策略是利用更多爬取的视频和更多计算资源进行训练,我们认为该策略效率低下:扩展世界模型还需要能够提供接地奖励信号的递归数据引擎。代码智能体的成功说明了这一点的重要性:由于代码可执行,编译器和运行时可为大语言模型(LLM)的强化学习(RL)后训练提供高质量奖励。相比之下,空间生成仍在很大程度上依赖CLIP分数等模糊代理,这些信号模糊且存在偏差,难以支持RL后训练。与之相比,游戏开发为空间世界模型提供了缺失的奖励环境:游戏引擎编码的场景是可执行的世界规范,引擎可高效检查碰撞、物理、可导航性和有限可玩性,而开发者则通过判断场景是否应被接受提供全局验证信号。游戏开发还为RL后训练提供了真实世界的长程轨迹数据。因此,我们提出了人类-引擎验证强化学习(RLHEV),这是一种结合了密集引擎信号与开发过程中隐含人类接受反馈的后训练范式。
英文摘要
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.