发表机构
Xiaomi EV; Northeastern University(小米汽车; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RoGe框架,通过端到端统一的隐式重建与生成实现新视角合成,在DL3DV数据集上,其性能优于各类基线,且联合训练及射线查询隐式特征均能提升效果。
AI 中文摘要
从稀疏输入进行新视角合成,既需要来自观测视角的几何基础,也需要未观测区域的生成先验,这推动了近期结合重建与生成的混合方法的发展。然而,现有方法通过渲染图像或点图、3D高斯等显式3D表示来衔接两者,因此生成过程依赖于场景有损且不完美的投影,继承了其误差,且重建无法从生成中获得修正信号。本文提出RoGe,一种端到端统一的重建与生成框架,移除了该显式衔接,其目标是在稀疏视角锚定的场景内漫游:给定少量带位姿的图像和一条相机轨迹,它沿该轨迹合成时间上连贯的视频。从稀疏输入视角出发,RoGe通过前馈重建模型构建隐式场景表示,并用目标相机射线查询该表示以获取逐视角几何特征,将这些特征作为条件注入视频扩散模型,无任何3D中间表示。两个模块联合训练,因此生成目标可直接塑造自身的几何条件。我们在DL3DV数据集上开展实验,结果显示RoGe在图像级指标和视频级时间一致性上均优于基于重建、基于生成及混合基线。 ablation研究证实,射线查询的隐式特征作为条件,其性能优于原始重建标记和渲染RGB,且联合训练可带来进一步提升。
英文摘要
Novel view synthesis from sparse inputs requires both geometric grounding from the observed views and generative priors of unobserved regions, motivating recent hybrid methods that combine reconstruction and generation. However, existing methods bridge the two with rendered images or explicit 3D representations such as point maps or 3D Gaussians. Generation is thus conditioned on a lossy and imperfect projection of the scene, inheriting its errors, and reconstruction receives no signal from generation to correct them. We present RoGe, an end-to-end unified reconstruction and generation framework that removes this explicit bridge. It targets roaming within a scene anchored by sparse views: given a few posed images and a camera trajectory, it synthesizes a temporally coherent video along that trajectory. From the sparse input views, RoGe builds an implicit scene representation with a feed-forward reconstruction model, and queries it with camera rays to obtain per-view geometric features. These features are injected into a video diffusion model as conditioning, without any explicit 3D intermediate. Both modules are trained jointly, so the generation objective directly shapes its own geometric conditioning. We conduct extensive experiments, where RoGe outperforms reconstruction-based, generation-based, and hybrid baselines in terms of image-level quality and video-level temporal and geometric consistency. Ablations confirm that ray-queried implicit features outperform both raw reconstruction tokens and rendered RGB as conditioning, and that joint training brings further gains. Our code will be released on https://jerry-locker.github.io/roge/.