发表机构
Westlake AGI Lab(西湖AGI实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Code World Model框架,以编码智能体为世界大脑生成可执行代码维持世界状态,结合视频模型实现视觉渲染,在游戏数据微调后取得良好效果,为开放式世界模型提供新路径。
AI 中文摘要
世界模型旨在模拟复杂环境在动作与事件作用下的演化过程,但现有基于视频的世界模型主要从视觉观测中学习动态,而视觉观测仅展现结果,未揭示支配世界演化的底层知识、规则与机制,这使得模型难以维持持久后果,也无法支持连贯、开放式的演化。本文提出Code World Model(代码世界模型),该框架通过结合语言模型的推理与编码能力,以及视频模型的生成先验,将世界演化与视觉实现分离开来。其中,编码智能体作为世界大脑,对事件及其后果进行推理,并生成可执行代码以维持持久的世界状态,执行符合规则的演化。为将可执行状态与视觉生成相连接,本文引入代理表示,该表示编码逐帧时空约束,并被编译为代理视频,用于条件驱动视频模型渲染高保真视觉观测。本文还开发了数据管道,用于从游戏和真实世界视频中构建对齐的代理-观测对。在基于配对游戏数据进行微调后,MiniMax-H3能够遵循编码智能体构建的简单交互世界中基于代理的时空规范,同时保留丰富的视觉细节与动态。这些结果证明了将用于持久世界演化的代码与用于灵活视觉实现的视频模型相结合的潜力,为开放式世界模型提供了新的路径。
英文摘要
World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics from visual observations, which reveal outcomes rather than the underlying knowledge, rules, and mechanisms governing world evolution. This makes it difficult to maintain persistent consequences and support coherent, open-ended evolution. We introduce Code World Model, a framework that separates world evolution from visual realization by combining the reasoning and coding capabilities of language models with the generative priors of video models. A coding agent serves as the world brain, reasoning about events and their consequences and generating executable code to maintain persistent world state and perform rule-consistent evolution. To connect executable state with visual generation, we introduce a proxy representation that encodes frame-wise spatiotemporal constraints and is compiled into a proxy video, which conditions a video model to render high-fidelity visual observations. We further develop data pipelines for constructing aligned proxy-observation pairs from gameplay and real-world videos. After fine-tuning on paired gameplay data, MiniMax-H3 follows proxy-based spatiotemporal specifications from simple interactive worlds built by the coding agent while preserving rich visual details and dynamics. These results demonstrate the potential of combining code for persistent world evolution with video models for flexible visual realization, providing a new path toward open-ended world models.
CommentsProject Page: https://buaacyw.github.io/cwm/