arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Genie Sim PanoWorld:基于全景场景建模与仿真的无限室内3D世界生成流水线

Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation

Yongxin Su, Linjie Hou, Feng Wang, Jialin Tang, Zhijun Li, Qian Wang, Maoqing Yao

arXiv 2607.26646首次发表:更新:

AI 中文总结

Genie Sim PanoWorld是两阶段前馈流水线,从单张360°全景生成可自由导航的高保真3D高斯场景,优于几何条件基线,且能零样本泛化至未见室内场景

AI 中文摘要

我们解决的问题是:无需针对每个场景进行优化,也无需多视角采集,仅从单个360°全景图重建高保真、可自由导航的3D场景。现有方法要么缺乏度量轨迹控制,阻碍了可靠的下游3D重建;要么在长距离相机运动下难以处理大遮挡,且需要高端多GPU this http URL。我们提出Genie Sim PanoWorld,这是一个两阶段前馈流水线,通过显式、轨迹可控的全景视频连接生成与重建。将由NavMesh规划的SE(3)漫游轨迹,通过密集几何扭曲的条件注入潜在视频扩散模型;长短轨迹混合训练与基于捷径模型的自一致性目标,共同在4步无CFG去噪中生成高保真视频。随后,一个前馈全景重建器将生成的视频提升为高保真3D高斯场景,支持实时自由视点漫游,可直接作为具身AI应用的仿真就绪资产。实验表明,Genie Sim PanoWorld在全景视频生成和下游3D重建上均优于几何条件基线,且对未见过的室内场景具备零样本泛化能力。

英文摘要

We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single $360^\circ$ panorama, without per-scene optimization or multi-view capture. Existing methods either lack metric trajectory control, which hinders reliable downstream 3D reconstruction, or struggle with large disocclusions under long-range camera motion while requiring high-end multi-GPU servers.We present Genie Sim PanoWorld, a two-stage feed-forward pipeline that bridges generation and reconstruction via an explicit, trajectory-controllable panoramic video. A NavMesh-planned $\mathrm{SE}(3)$ roaming trajectory is injected into a latent video diffusion model through dense geometry-warped conditioning; long--short trajectory mixed training and a self-consistency objective based on shortcut models together yield high-fidelity video in four CFG-free denoising steps. A feed-forward panoramic reconstructor then lifts the generated video into a high-fidelity 3D Gaussian scene that supports real-time, free-viewpoint roaming and can be directly used as a simulation-ready asset for embodied AI applications. Experiments show that Genie Sim PanoWorld outperforms geometry-conditioned baselines in both panoramic video generation and downstream 3D reconstruction, while generalizing zero-shot to unseen indoor scenes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑