GaussianSelector:基于图优化的三维高斯溅射中轻量型人引导物体选择方法
GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization
浏览论文内容
中文总结 AI 辅助
GaussianSelector是无需训练的轻量型框架,通过图优化从稀疏视点与涂鸦引导中选择三维物体,质量媲美多视图SAM方法,交互视点少、计算开销低,适用于三维场景编辑与资产提取。
中文摘要 AI 辅助
以最少的用户操作从重建场景中选择完整的三维物体,对于实际的场景编辑和具身交互至关重要。现有的基于3DGS(三维高斯溅射)的方法要么重新训练高斯表示以嵌入每个物体的标签,要么构建密集的多视图SAM( Segment Anything Model,图像分割模型)观测,两者都需要大量计算且需要密集的视点覆盖,而这在实际场景中很少能获得。我们提出了GaussianSelector,一种无需训练的框架,用于从稀疏视点和稀疏涂鸦引导中进行交互式三维物体选择。该方法直接在原生高斯基元上操作,将密集高斯粗化为几何连贯的超点,并利用外观和空间线索构建连续性加权图。通过感知可见性的透射覆盖将稀疏用户涂鸦提升到三维空间,并将选择问题求解为全局图割能量最小化,从而将稀疏证据传播到完整的三维物体。该设计天然支持多轮优化,用户可从额外视点迭代修正选择结果,逐步提升效果。实验表明,GaussianSelector在选择质量上与最先进的多视图SAM方法相当,同时需要显著更少的交互视点和低得多的计算开销。这些特性使其非常适合实际部署场景中的人在环三维场景编辑和三维资产提取。
英文摘要
Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene editing and embodied interaction. Existing 3DGS-based methods either retrain the Gaussian representation to embed per-object labels, or build dense multi-view SAM observations, both requiring heavy computation and dense viewpoint coverage that is rarely available in practice. We present GaussianSelector, a training-free framework for interactive 3D object selection from sparse views and sparse scribble guidance. Operating directly on native Gaussian primitives, we coarsen dense Gaussians into geometrically coherent superpoints and construct a continuity-weighted graph using appearance and spatial cues. Sparse user scribbles are lifted into 3D via visibility-aware transmittance coverage, and selection is solved as a global graph-cut energy minimization that propagates sparse evidence to a complete 3D object. This design naturally supports multi-round refinement, where users iteratively correct the selection from additional viewpoints to progressively improve the result. Experiments demonstrate that GaussianSelector achieves competitive selection quality against state-of-the-art multi-view SAM-based methods, while requiring significantly fewer interaction views and substantially lower computational overhead. These properties make it well suited for human-in-the-loop 3D scene editing and 3D asset extraction in real-world deployment scenarios.