arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12314cs.CV

StateFlow:用于预可视化的3D世界状态的构建、演化与访问

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

  • Beijing Jiaotong University(北京交通大学)
  • Mootion AI
  • Beijing Normal University(北京师范大学)
  • Peking University(北京大学)
  • Beijing Academy of Artificial Intelligence(北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao, Junnan Liu, Weirong Huang, Mengyu Wang, Tianxiao Fu, Yikai Wang, Peng-Shuai Wang, Xiaojie Jin, … 展开作者

Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao, Junnan Liu, Weirong Huang, Mengyu Wang, Tianxiao Fu, Yikai Wang, Peng-Shuai Wang, Xiaojie Jin, Yao Zhao, Yunchao Wei

AI总结:

StateFlow是一种以状态为中心的预可视化框架,通过构建、演化、访问可编辑3D世界的三阶段方法,提升生成预可视化的可控性与迭代编辑能力,可生成高质量3D世界用于视频创作和游戏原型设计。

AI中文摘要:

预可视化是电影、游戏、建筑和城市设计中创意与制作之间的中间层,它让创作者能够迭代地优化场景、动作、相机和时空动态。然而现有的生成方法依赖简单提示,通过单次图像或视频合成来共同控制所有这些因素,可控性弱且对迭代编辑的支持有限。从根本上说,世界包含多个具有几何、外观及其他属性的元素,再加上相机,不同帧通过对这个共享状态的局部修改或重组生成,否则该状态会被大量重复使用。因此,我们认为缺失的组件是一个显式且持久的工作状态。为解决此问题,我们提出StateFlow,一种以状态为中心的生成预可视化框架。StateFlow并非单次生成视频,而是使用可编辑的3D世界来组织场景结构、演化和相机,当需要更高保真度时,还可利用现成的视频模型提升视觉质量。该世界作为场景元素和相机配置的持久结构化3D状态被维护,是预可视化的核心工作表示。基于这一见解,StateFlow分为三个阶段来构建、演化和访问世界状态:状态构建阶段通过先验引导、感知冲突的双视图初始化,将生成的2D内容提升为连贯的3D世界;状态演化阶段将用户意图转化为结构化的状态转换,同时保留世界记忆,避免每次编辑都重新生成整个场景;状态访问阶段使用渲染反馈反射,将相机计划细化为视觉上可行的轨迹,避免仅依赖VLM语义。实验表明,StateFlow能生成用于视频创作和游戏类原型设计的高质量3D世界。

英文摘要:

Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors through one-shot image or video synthesis, offering weak controllability and limited support for iterative editing. Fundamentally, a world comprises multiple elements with geometry, appearance, and other attributes, together with cameras. Different frames are produced through local modifications or recombinations of this shared state, which is otherwise largely reused. Therefore, we argue that the missing component is an explicit and persistent working state. To address this, we present StateFlow, a state-centric framework for generative previsualization. Rather than generating videos in one shot, StateFlow uses an editable 3D world to organize scene structure, evolution, and cameras, while off-the-shelf video models enhance visual quality when higher fidelity is desired. This world is maintained as a persistent structured 3D state of scene elements and camera configurations, serving as the core working representation for previsualization. Built on this insight, StateFlow has three stages to construct, evolve, and access the world state. State construction lifts generated 2D content into a coherent 3D world through prior-guided, conflict-aware dual-view initialization, while State evolution translates user intent into structured state transitions while preserving world memory, avoiding full-scene regeneration for each edit. State access uses render-feedback reflection to refine camera plans into visually feasible trajectories, avoiding reliance on VLM semantics alone. Experiments show that StateFlow produces high-quality 3D worlds for video creation and game-like prototyping.

补充信息

↑