arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpatialCrafter:基于生成式3D代理的单图像世界建模

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan

arXiv 2608.27073首次发表:更新:

发表机构

Hong Kong University of Science and Technology; Tongyi Lab, Alibaba Group; ManyCore Tech Inc.; Jilin University(香港科技大学; 阿里巴巴集团通义实验室; ManyCore科技公司; 吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SpatialCrafter是解决图像到场景生成问题的两阶段框架,通过3D代理及相关策略提升3D一致性,构建了115K场景的新数据集,性能优于现有方法且鲁棒性强。

AI 中文摘要

可探索的图像到场景生成是游戏、机器人和虚拟现实等应用的核心需求。现有基于视频扩散模型(VDM)的方法通常依赖稀疏点云或2D全景图等不完整的条件信号,导致随机幻觉、长期漂移和3D一致性不佳等问题。我们提出SpatialCrafter,一种新颖的两阶段框架,通过引入全局3D代理解决上述问题以实现高保真的图像到场景生成。具体而言,我们将生成过程分解为全局代理生成与外观细化两个阶段:在代理生成阶段,我们提出Point锚定稀疏结构(PaSS)Flow模块,用于预测空间对齐且几何一致的3D代理;在外观细化阶段,我们将VDM重新定义为生成式延迟细化器,用于基于代理定义的场景几何合成高频逼真细节。为更好地将代理与预训练VDM集成,我们引入并行几何注入和代理感知损坏训练策略,在不破坏预训练生成流形的前提下提升对代理瑕疵的鲁棒性。此外,由于该可探索场景生成任务缺乏合适的数据集,我们构建了包含115K个场景的新大规模数据集,据我们所知,这是首个用于图像到场景生成的混合数据集。在合成数据集和真实世界数据集上的大量实验表明,SpatialCrafter的性能优于现有最先进方法,可缓解长期漂移问题,且在快速相机运动和极端视角变化下仍保持鲁棒性和一致性。代码、模型及新构建的数据集将公开发布,详情见此https URL。

英文摘要

Explorable image-to-scene generation is essential for applications in gaming, robotics, and virtual reality. Existing methods based on video diffusion model (VDM) commonly rely on incomplete conditioning signals such as sparse point clouds or 2D panoramas, leading to stochastic hallucinations, long-term drifts and suboptimal 3D consistency. We present SpatialCrafter, a novel two-stage framework that addresses these issues by introducing a global 3D proxy for high-fidelity image-to-scene generation. Specifically, we decompose the generation process into global proxy generation and appearance refinement. For proxy generation, we propose a Point-anchored Sparse Structure~(PaSS) Flow module that predicts a spatially aligned and geometrically consistent 3D proxy. For appearance refinement, we re-frame the VDM as a Generative Deferred Refiner which synthesizes high-frequency photorealistic details upon proxy-defined scene geometry. To better integrate the proxy with the pre-trained VDM, we introduce Parallel Geometry Injection and Proxy-Aware Corruption training strategies, which improve robustness to proxy artifacts without disrupting the pretrained generative manifold. Furthermore, as no suitable dataset exists for this explorable scene generation task, we construct a new large-scale dataset of 115K scenes. To the best of our knowledge, it is the first hybrid dataset for image-to-scene generation. Extensive experiments on both synthetic and real-world datasets show that SpatialCrafter outperforms state-of-the-art methods, mitigates long-term drift, and remains robust and consistent under rapid camera motion and extreme viewpoint changes. Our project page: \href{https://fangchuan.github.io/SpatialCrafter/}{fangchuan.github.io/SpatialCrafter/}

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑