arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27850cs.CV

GaussianDS:面向场景理解的深度监督语义高斯泼溅

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

Yufei Zhang, Chenlu Zhan, Hongwei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

GaussianDS提出深度监督的语义高斯泼溅框架,通过联合优化RGB、深度和语义并利用SAM2传播掩码,缓解语义漂移,在LERF和3D-OVS上取得SOTA性能。

中文摘要 AI 辅助

三维高斯泼溅(3D Gaussian Splatting)为三维重建提供了一种高效表示,近期扩展将语义属性附加到高斯上,用于开放词汇场景理解。然而,将依赖视角的二维基础模型输出提升到三维空间会引入跨视角不一致和几何基础薄弱的问题,导致严重的语义漂移和边界泄漏。我们提出GaussianDS,一种深度监督的语义三维高斯泼溅框架,将语义提升视为监督对齐问题,并从零开始联合优化RGB外观、渲染深度和紧凑语义。具体而言,GaussianDS将无序多视角图像组织成姿态感知的伪视频轨迹,通过SAM2传播视角一致的掩码。在联合优化过程中,尺度-平移对齐的单目深度监督和深度全变分正则化稳定高斯几何,而深度边缘感知的细化损失将语义过渡显式锚定到物理几何不连续处。大量评估表明,我们的端到端框架不仅保留了高保真三维重建和实时渲染,还建立了优越的语义理解。GaussianDS通过缓解语义泄漏,在LERF(60.5% mIoU)和3D-OVS(97.79% mIoU,90.28% mBIoU)上取得了新的最先进性能,同时无缝促进下游三维物体移除。

英文摘要

3D Gaussian Splatting provides an efficient representation for 3D reconstruction, and recent extensions attach semantic attributes to Gaussians for open-vocabulary scene understanding. However, lifting view-dependent 2D foundation-model outputs into 3D space introduces cross-view inconsistencies and weak geometric grounding, leading to severe semantic drift and boundary leakage. We propose GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch. Specifically, GaussianDS organizes unordered multi-view images into a pose-aware pseudo-video trajectory to propagate view-consistent masks via SAM2. During joint optimization, scale-shift-aligned monocular depth supervision and depth total-variation regularization stabilize Gaussian geometry, while a depth-edge-aware refinement loss explicitly anchors semantic transitions onto physical geometric discontinuities. Extensive evaluations show that our end-to-end framework not only retains high-fidelity 3D reconstruction and real-time rendering, but also establishes superior semantic understanding. GaussianDS sets new state-of-the-art performance on LERF (60.5% mIoU) and 3D-OVS (97.79% mIoU, 90.28% mBIoU) by mitigating semantic leakage, while seamlessly facilitating downstream 3D object removal.

发表机构

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

↑