可编程世界模型
Programmable World Model
浏览论文内容
中文总结 AI 辅助
提出可编程世界模型,将状态演化与视觉生成解耦,通过程序控制实体状态,在CombatStateBench上实现94%计数和98%状态准确率,支持持久可编程世界。
中文摘要 AI 辅助
近期视频世界模型生成的视觉体验日益逼真且具有交互性,但缺乏可靠的机制来维持持久的世界状态,并在长时间交互中强制执行可编程规则。我们提出了可编程世界模型(Programmable World Model),这是一种将世界状态演化与视觉观察生成解耦的框架。智能体将自然语言指令翻译为可执行程序,这些程序指定实体状态和状态转换规则,从而能够直接控制单个实体及其交互。一个轻量级引擎执行这些程序,以更新并维护显式、持久化的全局世界状态,包括屏幕外实体和非视觉属性。为了将世界状态与视觉生成相连接,我们引入了状态增强的三维有向边界框(OBBs)作为中间表示。该表示连同目标相机轨迹,被确定性地编译为像素对齐的时空条件信号,用于作为生成式渲染器的预训练视频模型。这一设计允许用户创建具有预定义机制的可玩游戏,直接控制单个实体,并在整个游戏过程中保持持久的世界状态。我们进一步引入了CombatStateBench,一个用于评估可编程世界模型的基准。在CombatStateBench上,我们的方法实现了94%的计数准确率和98%的状态准确率,大幅优于现有的交互式视频世界模型,同时支持连贯的长时程生成。这些结果证明了将显式状态演化与生成式渲染分离对于构建持久、可编程世界的有效性。
英文摘要
Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visual observation generation. An agent translates natural-language instructions into executable programs that specify entity states and state-transition rules, enabling direct control over individual entities and their interactions. A lightweight engine executes these programs to update and maintain an explicit, persistent global world state, including off-screen entities and non-visual attributes. To connect world state with visual generation, we introduce state-augmented 3D oriented bounding boxes (OBBs) as an intermediate representation. This representation, together with the target camera trajectory, is deterministically compiled into pixel-aligned spatiotemporal conditioning signals for a pretrained video model serving as the generative renderer. This design allows users to create playable games with predefined mechanics, direct control over individual entities, and persistent world state throughout gameplay. We further introduce CombatStateBench, a benchmark for evaluating programmable world models. On CombatStateBench, our method achieves 94% Count Accuracy and 98% State Accuracy, substantially outperforming existing interactive video world models while supporting coherent long-horizon generation. These results demonstrate the effectiveness of separating explicit state evolution from generative rendering for building persistent, programmable worlds.
发表机构
- Alaya Lab(Alaya实验室)
机构由 AI 辅助整理,请以论文原文为准。