AI 中文总结
Genie Sim PanoWorld是两阶段前馈流水线,从单张360°全景生成可自由导航的高保真3D高斯场景,优于几何条件基线,且能零样本泛化至未见室内场景
AI 中文摘要
我们解决的问题是:无需针对每个场景进行优化,也无需多视角采集,仅从单个360°全景图重建高保真、可自由导航的3D场景。现有方法要么缺乏度量轨迹控制,阻碍了可靠的下游3D重建;要么在长距离相机运动下难以处理大遮挡,且需要高端多GPU this http URL。我们提出Genie Sim PanoWorld,这是一个两阶段前馈流水线,通过显式、轨迹可控的全景视频连接生成与重建。将由NavMesh规划的SE(3)漫游轨迹,通过密集几何扭曲的条件注入潜在视频扩散模型;长短轨迹混合训练与基于捷径模型的自一致性目标,共同在4步无CFG去噪中生成高保真视频。随后,一个前馈全景重建器将生成的视频提升为高保真3D高斯场景,支持实时自由视点漫游,可直接作为具身AI应用的仿真就绪资产。实验表明,Genie Sim PanoWorld在全景视频生成和下游3D重建上均优于几何条件基线,且对未见过的室内场景具备零样本泛化能力。
英文摘要
We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single $360^\circ$ panorama, without per-scene optimization or multi-view capture. Existing methods either lack metric trajectory control, which hinders reliable downstream 3D reconstruction, or struggle with large disocclusions under long-range camera motion while requiring high-end multi-GPU servers.We present Genie Sim PanoWorld, a two-stage feed-forward pipeline that bridges generation and reconstruction via an explicit, trajectory-controllable panoramic video. A NavMesh-planned $\mathrm{SE}(3)$ roaming trajectory is injected into a latent video diffusion model through dense geometry-warped conditioning; long--short trajectory mixed training and a self-consistency objective based on shortcut models together yield high-fidelity video in four CFG-free denoising steps. A feed-forward panoramic reconstructor then lifts the generated video into a high-fidelity 3D Gaussian scene that supports real-time, free-viewpoint roaming and can be directly used as a simulation-ready asset for embodied AI applications. Experiments show that Genie Sim PanoWorld outperforms geometry-conditioned baselines in both panoramic video generation and downstream 3D reconstruction, while generalizing zero-shot to unseen indoor scenes.