发表机构
USTC; TongYi Lab(中国科学技术大学; 通义实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CoDeR范式,用代码显式构建可执行世界,结合视频生成实现视觉化,支持长期记忆、开放式交互与多智能体演化,性能达到最先进水平。
AI 中文摘要
我们引入了CoDeR,一种用于世界建模的新范式。与现有通过视觉观察隐式表示世界动态的视频世界模型不同,我们的系统通过代码显式构建一个可执行的世界,并采用视频生成模型进行视觉实现。具体而言,我们协调五个互补角色,将高层概念转化为结构化的世界规则、可执行的动态和感知观察。这种设计实现了长期记忆、开放式交互、自主世界演化以及多智能体场景,其中多个实体可以持续行动、交互和演化,超越当前观察范围。大量实验表明,我们的框架显著扩展了现有世界模型的能力,实现了长期记忆、开放式交互、自主演化和持久的多智能体动态,同时在多个评估设置中达到了最先进的性能。代码和模型权重将公开提供。项目页面:CoDeR。
英文摘要
We introduce \textbf{CoDeR}, a new paradigm for world modeling. Unlike existing video world models that implicitly represent world dynamics through visual observations, our system explicitly constructs an executable world with code and employs video generation models for visual realization. Specifically, we coordinate five complementary roles to translate high-level concepts into structured world rules, executable dynamics, and perceptual observations. This design enables \textit{long-term memory}, \textit{open-ended interactions}, \textit{autonomous world evolution}, and \textit{multi-agent scenarios}, where multiple entities can act, interact, and evolve persistently beyond the current observation. Extensive experiments demonstrate that our framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings. Code and model weights will be made publicly available. Project Page: \href{https://becauseimbatman0.github.io/CoDeR}{CoDeR}.
Commentshttps://becauseimbatman0.github.io/CoDeR