Z-Order Transformer用于前馈高斯点云
Z-Order Transformer for Feed-Forward Gaussian Splatting
- The University of Hong Kong(香港大学)
- Futurewei Technologies Inc(未来科技公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Z-Order Transformer用于前馈高斯点云,通过稀疏注意力机制和Z-order策略有效捕捉高斯点间空间与语义关系,提升实时渲染质量与效率。
AI中文摘要:
近年来,3D高斯点云(3DGS)在逼真新视角合成中取得显著进展。然而,传统3DGS依赖缓慢的迭代优化过程,限制了实时场景的应用。为克服这一瓶颈,近期的前馈方法试图直接从图像预测高斯属性,但常面临高斯原始体冗余和渲染质量的挑战。本文引入基于Transformer的架构专门用于前馈高斯点云。我们的核心思想是通过稀疏注意力机制捕捉高斯点间的空间和语义关系,该机制由Z-order策略实现,将无序的高斯点集组织成空间连贯的序列。此外,我们采用此Z-order策略来适应性地抑制冗余,同时保留关键结构细节。这使Transformer能够高效建模上下文,压缩高斯原始体,并在单次前向传递中预测高斯属性。全面实验表明,我们的方法在较少高斯原始体的情况下实现了快速且高质量的新视角合成。
英文摘要:
Recent advances in 3D Gaussian Splatting (3DGS) have enabled significant progress in photorealistic novel view synthesis. However, traditional 3DGS relies on a slow, iterative optimization process, which limits its use in scenarios demanding real-time results. To overcome this bottleneck, recent feed-forward methods aim to predict Gaussian attributes directly from images, but they often struggle with the redundancy of Gaussian primitives and rendering quality. In this work, we introduce a transformer-based architecture specifically designed for feed-forward Gaussian Splatting. Our key insight is that spatial and semantic relationships among Gaussians can be effectively captured through a sparse attention mechanism, enabled by a Z-order strategy that organizes the unstructured Gaussian set into a spatially coherent sequence. Furthermore, we incorporate this Z-order strategy to adaptively suppress redundancy while preserving critical structural details. This allows the transformer to efficiently model context, compress Gaussian primitives, and predict Gaussian attributes in a single forward pass. Comprehensive experiments demonstrate that our method achieves fast and high-quality novel view synthesis with fewer Gaussian primitives.