arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StereoSplat+:具有扩散辅助渐进推理的前馈立体高斯点云渲染

StereoSplat+: Feed-Forward Stereo Gaussian Splatting with Diffusion-Assisted Progressive Inference

Zihua Liu, Masatoshi Okutomi

arXiv 2607.08808首次发表:更新:

发表机构

Institute of Science Tokyo(东京科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究从单目立体观察恢复高质量3DGS场景的问题,提出StereoSplat+框架,包含StereoSplat估计器及扩散增强单步渐进推理方案,实验表明该方法提升了新视图渲染质量和几何精度,优于现有前馈3DGS基线。

AI 中文摘要

3D高斯点云渲染的最新进展为新视图合成提供了高质量、可渲染的场景表示。然而,大多数现有3DGS管道依赖多视图观察(或对未来帧的非因果访问)来实现足够的覆盖,这在设备上的机器人技术和AR设置中通常不可用,因为传感仅限于单个立体相机。从单目立体观察中恢复高质量的3DGS场景仍然具有挑战性。我们提出了StereoSplat+,这是一个扩散增强的前馈框架,能够从单个立体对进行因果重建。我们的方法基于两个关键组件。首先,我们提出了StereoSplat,这是一种输入不变的前馈3D高斯估计器,它将可变数量的姿态立体对作为输入,并预测高质量的3D高斯。StereoSplat通过成本体积分支和基于三平面的3D体积分支融合互补几何线索,并利用连续姿态编码在视图数量和相机配置之间进行泛化。其次,由于在推理时通常无法获得多个姿态立体对,我们引入了一种称为StereoSplat+的扩散增强单步渐进推理方案:从一个立体对开始,我们从预测的3DGS渲染新的立体视图,用单步扩散增强器对其进行细化,并将它们作为额外输入反馈以更新3DGS。在KITTI-360数据集上的实验表明,StereoSplat+提高了新视图渲染质量和几何精度,特别是在遮挡区域和强视图外推下,优于最近的前馈3DGS基线。

英文摘要

Recent advances in 3D Gaussian Splatting (3DGS) have enabled high-quality, render-ready scene representations for novel-view synthesis. However, most existing 3DGS pipelines rely on multi-view observations (or non-causal access to future frames) to achieve sufficient coverage, which is often unavailable in on-device robotics and AR settings where sensing is restricted to a single stereo rig. Recovering a high-quality 3DGS scene from one stereo observation, therefore, remains challenging due to occlusions, limited field of view, and missing geometry. We present StereoSplat+, a diffusion-enhanced feed-forward framework that enables causal reconstruction from a single stereo pair. Our method builds on two key components. First, we propose StereoSplat, an input-invariant feed-forward 3D Gaussian estimator that takes a variable number of posed stereo pairs as input and predicts high-quality 3D Gaussians. StereoSplat fuses complementary geometry cues via a cost-volume branch and a triplane-based 3D volume branch and leverages continuous pose encoding to generalize across view counts and camera configurations. Second, since multiple posed stereo pairs are typically unavailable at inference time, we introduce a diffusion-enhanced one-shot progressive inference scheme called StereoSplat+: starting from one stereo pair, we render novel stereo views from the predicted 3DGS, refine them with a one-step diffusion enhancer, and feed them back as additional inputs to update the 3DGS. Experiments on the KITTI-360 dataset show that StereoSplat+ improves novel-view rendering quality and geometry accuracy, especially in occluded regions and under strong view extrapolation, outperforming recent feed-forward 3DGS baselines.

Comments8 pages, accepted as a conference paper for IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑