arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16863cs.CV

SplatGuide:用于无姿态新视图合成的3D高斯几何先验

SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis

Yejun Zhang, Zihan Wang, Xu Ji, Yihao Wang, Yuxin Hou, Junyuan Fang, Juho-Matti Kilpeläinen, Arno Solin, Hamed Rezazadegan Tavakoli, Esa Rahtu, Juho Kannala

首次发表
浏览论文内容

中文总结 AI 辅助

SplatGuide复用单个3DGS场景生成三种互补信号,实现无姿态新视图合成,在多个基准数据集上达到SOTA,在RealEstate10K上超越真实姿态基线。

中文摘要 AI 辅助

从无姿态图像生成逼真的新视图,既需要3D几何理解能力,也需要合成未见内容的能力。一种自然的策略是将前馈3DGS重建与多视图扩散相结合,但现有流水线最多从重建中提取一种信号,要么是像素渲染,要么是学习到的特征,没有利用每个高斯的可见性来进行感知遮挡的参考选择。这种“信息脱节”使得可渲染几何、可见性线索和学习到的特征都未被利用。SplatGuide通过在三个互补角色中复用同一个3DGS场景来弥合这一脱节:渲染图像提供像素对齐的几何条件,每个高斯的源视图索引被渲染为目标视图投票图以进行感知遮挡的参考选择,重建标记通过交叉注意力提供特征级指导,所有三种信号都来自同一重建前向传播。在RealEstate10K、DL3DV、Tanks-and-Temples和Mip-NeRF 360数据集上,SplatGuide实现了无姿态新视图合成的SOTA,在RealEstate10K上,当输入视图数量适中时,它超越了真实姿态基线。

英文摘要

Generating photorealistic novel views from unposed images requires both 3D geometric understanding and the ability to synthesize unseen content. A natural strategy combines feed-forward 3DGS reconstruction with multi-view diffusion. Yet prior pipelines extract at most one signal from the reconstruction, either pixel rendering or learned features, while none exploits per-Gaussian visibility for occlusion-aware reference selection. This *information disconnect* leaves renderable geometry, visibility cues, and learned features unused. SplatGuide closes this disconnect by reusing a single 3DGS scene across three complementary roles. Rendered images provide pixel-aligned geometric conditioning. Per-Gaussian source-view indices are rendered into a target-view voting map for occlusion-aware reference selection. Reconstruction tokens supply feature-level guidance via cross-attention. All three signals derive from the same reconstruction forward pass. Across RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360, SplatGuide achieves state-of-the-art pose-free novel view synthesis. On RealEstate10K, with a moderate number of input views, it surpasses the ground-truth-pose baseline.

发表机构

  • Aalto University(阿尔托大学)
  • Deep Render(深度渲染公司)
  • ELLIS Institute Finland(芬兰ELLIS研究所)
  • Nokia Technologies(诺基亚技术公司)
  • Tampere University(坦佩雷大学)
  • University of Oulu(奥卢大学)

机构由 AI 辅助整理,请以论文原文为准。

↑