arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07937cs.CV

FlexSplat:无需点云对应关系的灵活前馈三维高斯溅射方法

FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence

Amir Sabbaghziarani, Hanting Ye, Maria Gorlatova, Yi Ding

首次发表
浏览论文内容

中文总结 AI 辅助

FlexSplat是一种无需相机位姿的前馈三维高斯溅射框架,联合训练几何Transformer与高斯解码器,在ShapeNet-SRN和GSO数据集上实现了与已知位姿方法相当的新视图合成性能。

中文摘要 AI 辅助

我们提出了FlexSplat,一种从无标定、以物体为中心的多视图图像集合中生成新视图合成(NVS)的前馈框架。近期一系列基于查询的方法将三维高斯视为Transformer查询,通过多视图可变形注意力对其进行优化,以此重建紧凑的三维高斯集合,但这些方法均假设相机位姿是已知的。FlexSplat摒弃了这一假设:我们联合训练几何Transformer与高斯解码器,以预测每张图像的相机参数和深度,这些参数和深度进而构建深度引导的高斯参数化,以及多视图可变形交叉注意力,该注意力会聚合所有输入视图的证据,形成单一视图一致的基元集合。我们采用不确定性加权的深度一致性目标,使联合训练的几何模型适配重建任务;同时,解码过程中形成的跨视图一致性可吸收估计相机和深度的残差误差。该表示采用紧凑的高斯预算,且与输入分辨率解耦——不同于像素对齐方法,其基元数量不会随图像网格增长,也不依赖于视图数量。在ShapeNet-SRN和Google Scanned Objects(GSO)数据集上,FlexSplat的性能与采用已知相机位姿的最先进重建方法相当或接近,且无需相机位姿或真实深度;在GSO数据集上,其感知质量(LPIPS)与对比方法中的最优水平相当。我们的结果表明,联合训练的几何前端足以使基于查询的高斯重建实现无标定操作,同时与已知位姿方法的PSNR差距保持在0.7 dB以内,且感知质量相当。

英文摘要

We present FlexSplat, a feed-forward framework for novel view synthesis (NVS) from uncalibrated, object-centric multi-view image collections. A recent line of query-based methods reconstructs a compact set of 3D Gaussians by treating them as transformer queries that are refined with multi-view deformable attention; these methods, however, assume that camera poses are given. FlexSplat removes this assumption: a geometry transformer is trained jointly with the Gaussian decoder to predict per-image camera parameters and depth, which in turn ground a depth-guided Gaussian parameterization and a multi-view deformable cross-attention that aggregates evidence across all input views into a single, view-consistent set of primitives. An uncertainty-weighted depth-consistency objective lets the jointly trained geometry adapt to the reconstruction task, while the cross-view consensus formed during decoding absorbs the residual error of the estimated cameras and depth. The representation uses a compact Gaussian budget that is decoupled from the input resolution - unlike pixel-aligned methods, the primitive count does not grow with the image grid - and is not dictated by the number of views. On ShapeNet-SRN and Google Scanned Objects (GSO), FlexSplat matches or approaches posed state-of-the-art reconstructors while requiring neither camera poses nor ground-truth depth, and matches the best perceptual (LPIPS) quality among the compared methods on GSO. Our results indicate that a jointly trained geometry front-end is sufficient to bring calibration-free operation to query-based Gaussian reconstruction while staying within 0.7 dB PSNR of posed methods and matching their perceptual quality.

发表机构

  • Tri-Institutional Georgia Institute of Technology/Georgia State University/Emory University Center for Translational Research in Data Science and Neuroimaging (TReNDS)(三机构佐治亚理工学院/佐治亚州立大学/埃默里大学转化数据科学与神经影像中心(TReNDS))
  • Georgia State University(佐治亚州立大学)
  • Duke University(杜克大学)
  • University of Tennessee, Knoxville(田纳西大学诺克斯维尔分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑