S2Tok:基于持久空间令牌的流式3D高斯重建
S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens
浏览论文内容
中文总结 AI 辅助
S2Tok提出前馈框架,通过空间感知Transformer和准入模块维护持久场景令牌,实现无标定图像流的流式3D高斯重建,在四个基准上达到竞争性渲染质量。
中文摘要 AI 辅助
流式3D重建需要的不仅仅是一系列几何预测:它需要一个持久的场景状态,该状态能够整合新证据,并在观测到达时保持可渲染性。潜在空间令牌为此目的提供了一种有前景的表示,但从图像集合中构建它们留下了如何在线维护的问题,其中每次观测可能既重新访问已知区域又揭示新内容。我们提出了S2Tok,一个前馈框架,从无标定图像流中维护一个大小自适应、持久的场景状态。其核心思想是区分对现有表示的更新与选择性扩展。一个空间感知的Transformer将每个传入的观测与持久场景令牌整合,而一个学习到的准入模块选择性地扩展表示以限制冗余存储。一个分层解码器和高斯头将演化的状态转换为非像素对齐的3D高斯,从而实现无需缓存先前帧的新视角渲染。在四个基准上的实验展示了具有紧凑高斯表示的竞争性流式渲染质量。这些结果支持潜在空间令牌作为在线3D重建的持久计算状态,将学习到的场景更新与显式高斯渲染相结合。
英文摘要
Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection leaves open how to maintain them online, where each observation may both revisit known regions and reveal new content. We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams. Its central idea is to distinguish updates to the existing representation from selective expansion. A spatially informed transformer integrates each incoming observation with the persistent scene tokens, while a learned admission module selectively expands the representation to limit redundant storage. A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames. Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations. These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.