基于显著性引导基元合并的紧凑前馈三维高斯模型
Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging
查看机构详情
- Bosch Research(博世研究院)
- University of Bonn(波恩大学)
- Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对前馈三维高斯方法基元冗余问题,提出显著性引导的基元合并流水线,可在仅1/20高斯数量下保留视觉质量,实现高效紧凑的三维场景表示与渲染。
中文摘要 AI 辅助
三维场景重建、建模与渲染与众多任务高度相关,三维高斯溅射(3D Gaussian splatting)已成为该场景下的标准选择。其前馈变体可从稀疏输入视图快速重建,但常生成逐像素基元,导致表示高度冗余、效率低下。本文提出一种结构感知合并流水线,可将任意前馈方法的逐像素基元整合为紧凑、内容自适应的高斯集合,同时在仅为逐像素方法1/20数量的高斯下,基本保留视觉质量。该流水线通过显著性图引导的自适应超像素分割,将空间相干、外观相似的高斯分组为可变大小簇,在纹理区域分配精细片段,在均匀区域分配粗糙片段;通过学习编码器将每个簇压缩为紧凑潜在表示,再通过学习合并器基于几何重叠与特征相似性匹配并整合跨视图表示;最后通过细节层次解码器生成可控分辨率的最终高斯,支持推理时灵活的质量-效率权衡。作为后处理模块,该流水线与骨干网络无关,可利用现有前馈方法的优势,相比以往以减少基元数量为目标的方法,实现更优、更鲁棒的质量,同时提供可高效渲染的高度紧凑表示。
英文摘要
3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this context. Its feed-forward variants provide fast reconstruction from sparse input views but often produce per-pixel primitives, leading to highly redundant and thus inefficient representations. We present a structure-aware merging pipeline that takes per-pixel primitives from any feed-forward method and consolidates them into a compact, content-adaptive Gaussian set while largely retaining visual quality at just $\frac{1}{20}^\text{th}$ of the Gaussians of a per-pixel method. We group spatially coherent Gaussians of similar appearance into variable-size clusters via adaptive superpixel segmentation guided by a saliency map, which allocates fine segments to textured regions and coarse segments to homogeneous areas. We compress each cluster into a compact latent representation through a learned encoder, then match and consolidate representations across views based on geometric overlap and feature similarity via a learned merger. A level-of-detail decoder then produces the final Gaussians at a controllable resolution, enabling a flexible quality-efficiency trade-off at inference. As a post-processing module, the pipeline is backbone-agnostic, leveraging the strengths of existing feed-forward methods. This leads to better and more robust quality than achieved by previous approaches that target a reduction in primitive count, while providing a highly compact representation, that can be rendered efficiently.