OctWorld:基于八叉树三维映射的长程世界一致视频生成
OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping
浏览论文内容
中文总结 AI 辅助
研究提出带持久三维记忆的视频扩散框架OctWorld,通过引入空间自适应三维记忆OctMap解决长程生成的空间一致性难题,生成的视频在基准及长程设置下优于现有方法。
中文摘要 AI 辅助
我们提出OctWorld,这是一个带有持久三维记忆的视频扩散框架,用于生成可探索、世界一致且高保真的视觉场景。给定单张图像,OctWorld会沿用户指定的相机轨迹执行稳定的自回归世界生成。我们聚焦于长程生成,其特征是扩展的相机路径和宽视角覆盖,在此过程中,当重新访问先前生成的区域时,保持空间一致性极具挑战性。为解决该问题,我们引入OctMap,这是一种可扩展且空间自适应的三维记忆,会逐步将生成的视觉观测及其对应的深度图融合为全局表示。OctMap在动态稀疏八叉树内采用TSDF融合,其空间分辨率会根据图像证据进行调整。该设计可在不同场景尺度上保留几何和外观细节,同时维持低内存开销。实验表明,OctWorld可生成长程、空间一致的视频,在现有基准和具有挑战性的长程生成设置上均优于先前方法;OctMap相比基于点的缓存和固定分辨率TSDF体也具有明显优势。项目页面:this https URL
英文摘要
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-consistent, and high-fidelity visual scenes. Given a single image, OctWorld performs stable autoregressive world generation along user-specified camera trajectories. We focus on long-range generation, characterized by extended camera paths and wide viewpoint coverage, where preserving spatial consistency is particularly challenging when previously generated regions are revisited. To address this problem, we introduce OctMap, an extensible and spatially adaptive 3D memory that progressively fuses generated visual observations and their corresponding depth maps into a global representation. OctMap employs TSDF fusion within a dynamic sparse octree whose spatial resolution adapts to image evidence. This design preserves geometric and appearance details across diverse scene scales while maintaining low memory overhead. Experiments demonstrate that OctWorld generates long-range, spatially consistent videos and outperforms prior methods on both existing benchmarks and challenging long-range generation settings. OctMap also provides clear advantages over point-based caches and fixed-resolution TSDF volumes. Project page: https://maxtirerror.github.io/octworldpage/
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Microsoft Research Asia(微软亚洲研究院)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。