arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldWeave:为视频生成构建持续增长的持久几何世界

WorldWeave: Growing Persistent Geometric Worlds for Video Generation

Yifan Huang, Lifan Jiang, Qingyue Hao, Cheng Chen, Boxi Wu, Xiaoxue Ren, Xiaofei He, Dehai Zhao

arXiv 2609.34221首次发表:更新:

发表机构

Zhejiang University; Daerwen AI(浙江大学; 达尔文AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

WorldWeave通过将世界状态维护与视觉渲染解耦,结合持续高程图生成和智能体引导的场景组织,构建可扩展的显式3D世界,实现视频生成中的持久结构记忆与一致几何基础。

AI 中文摘要

尽管取得了快速进展,世界模型仍然缺乏显式的、持久的结构记忆,这使得在持续的场景扩展和跨视角重访过程中难以保持一致的世界结构。为解决这一局限,我们提出了WorldWeave,一个将世界状态维护与视觉渲染解耦的世界生成框架。具体而言,WorldWeave结合了持续的高程图生成与智能体引导的场景组织和拼接,以构建一个可扩展的显式3D世界,该世界在保留现有结构的同时逐步扩展结构记忆。首先,其地形模块使用基于扩散的图像外推,在邻域条件和边界约束下生成连续的度量高程图。接下来,一个智能体整合用户意图、地形证据和跨区域连通性约束,通过分层语义规划、确定性几何编译和局部修订来构建场景。最后,在视觉生成过程中,规划的相机轨迹通过只读接口查询世界几何,生成深度序列以指导视频合成,而不会将生成结果写回世界状态。因此,结构记忆独立于短窗口视频生成,使得无需预定义地图边界即可持续扩展,并为跨轨迹和重复访问的观测提供一致的几何基础。

英文摘要

Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.

CommentsProject page: https://laiyindagm.github.io/WorldWeave/ . Code repository: https://github.com/laiyindagm/WorldWeave (implementation coming soon)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑