arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01659cs.CV

StreamSplat:流式前馈三维高斯溅射

StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting

Changhao Song, Yuxuan Wang, Qibiao Li, Youcheng Cai, Ligang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出StreamSplat流式前馈3DGS框架,通过VACC、HPDA、CGFI技术实现长输入流下的高效因果场景更新,在多数据集上的新视角合成质量优于固定视角基准。

中文摘要 AI 辅助

前馈三维高斯溅射技术可实现高效的新视角合成,无需针对每个场景单独优化,但现有多数方法均假设存在固定的上下文视角集合,并对其进行联合处理。这一特性限制了它们在在线场景中的适用性,在线场景中校准后的视角会依次到达,且场景必须进行因果更新。本文提出StreamSplat,这是一种流式前馈三维高斯溅射(3DGS)框架,可逐步维护持久的、基于几何的场景状态,并在每个输入块后将其解码为可渲染的三维高斯。StreamSplat的核心是体素对齐因果缓存(VACC),该缓存以内存受限的体素结构存储历史三维令牌,因此内存会随已探索的场景几何增长,而非随流长度增长。为在因果预测中更好地复用历史信息,本文引入历史投影深度锚定(HPDA),将缓存的几何投影为当前代价体估计的深度引导;还引入缓存引导特征注入(CGFI),将缓存的潜在证据注入高斯令牌回归。在DL3DV、RealEstate10K和ScanNet数据集上的实验表明,StreamSplat在稀疏因果输入下仍与最先进的前馈3DGS方法具有竞争力,且无需使用未来视角或全场景上下文;更重要的是,它可扩展到包含256、512和1024个视角的长输入流,而固定视角基准会出现内存不足问题,随着更多观测的到来,其新视角合成质量持续提升。代码将在论文接收后公开。

英文摘要

Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set of context views and process them jointly. This limits their applicability to online scenarios where calibrated views arrive sequentially and the scene must be updated causally. We present \emph{StreamSplat}, a streaming feed-forward 3DGS framework that incrementally maintains a persistent geometry-grounded scene state and decodes it into renderable 3D Gaussians after each input chunk. StreamSplat centers on a \textbf{Voxel-Aligned Causal Cache (VACC)}, which stores historical 3D tokens in a memory-bounded voxel structure so that memory grows with explored scene geometry rather than stream length. To better reuse history during causal prediction, we introduce \textbf{History-Projected Depth Anchoring (HPDA)} to project cached geometry as depth guidance for current cost-volume estimation, and \textbf{Cache-Guided Feature Injection (CGFI)} to inject cached latent evidence into Gaussian-token regression. Experiments on DL3DV, RealEstate10K, and ScanNet show that StreamSplat remains competitive with state-of-the-art feed-forward 3DGS methods under sparse causal inputs, despite not using future views or full-scene context. More importantly, it scales to long input streams with 256, 512, and 1024 views where fixed-view baselines run out of memory, yielding sustained improvements in novel-view synthesis quality as more observations arrive. The code will be made publicly available upon acceptance.

↑