Oneira:从视频世界模型中的开放式生成到开放世界交互
Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models
浏览论文内容
中文总结 AI 辅助
Oneira通过编码智能体管理显式世界状态,实现视频世界模型中生成与交互的闭环,支持新对象纳入交互并持久化状态变化,实验验证了长时程交互一致性。
中文摘要 AI 辅助
生成式视频世界模型现在能够合成开放式的环境,智能体可以在其中以简单的方式进行导航和交互。然而,开放式生成并不意味着完全交互:随着生成世界的扩展,通过导航新创建的内容应该扩展智能体可以作用的对象,并且当智能体改变世界时,这些改变应该成为环境的持久部分,而不是短暂的视觉效果。我们将这两个要求分别定义为“开放世界交互性”,即新生成或遇到的对象被纳入可操作的世界中,以及“持久状态”,即交互结果被提交到世界状态,并持续影响后续的观察和交互。我们提出了Oneira,一个交互式视频世界模型,通过由编码智能体管理的显式、可扩展的世界状态,闭环了生成与交互之间的循环。给定当前观察和一个动作或高级目标,智能体读取世界状态,定位相关实体,规划交互,并将其结果写回世界状态表。当探索揭示新对象时,智能体从生成的观察中将其纳入,从而允许交互空间随生成世界扩展。同时,先前引起的状态变化会跨视频片段传递,使交互的后果成为后续世界演化的持久部分。更新后的世界状态沿着相机动作轨迹渲染成粗略的条件视频,视频生成器从中填充状态中未表示的外观、运动和交互细节。实验表明,Oneira能够直接且一致地与新生成的对象交互,同时在长时程内保留先前交互的效果。项目页面:此https URL
英文摘要
Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not represented in the state. Experiments show that Oneira enables direct and consistent interaction with newly generated objects, while preserving the effects of prior interactions over long horizons. Project page: https://madaoer.github.io/projects/oneira
发表机构
- Monash University(莫纳什大学)
- Dalian University of Technology(大连理工大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Oxford University(牛津大学)
- University of Bristol(布里斯托大学)
机构由 AI 辅助整理,请以论文原文为准。