混合高斯体用于鲁棒开放词汇3D分割:多视角对象关联与边界细化
Hybrid Gaussians for Robust Open-Vocabulary 3D Segmentation with Multi-View Object Association and Boundary Refinement
- Durham University(杜伦大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对开放词汇3D分割中多视角身份不稳定和对象判别性弱的问题,提出混合高斯体表示,结合多视角对象关联与边界重建优化,在LERF上达到59.1% mIoU,相对提升13.4%。
AI中文摘要:
开放词汇3D分割能够根据自由形式的文本查询定位对象,但在真实图像序列中仍面临挑战:不完整或含噪的2D监督会破坏多视角身份分配的稳定性,而全场景语义学习则削弱了对象级别的判别能力。我们提出了混合高斯体(Hybrid Gaussians),一种统一的三维表示,联合建模对象关联与语言对齐的语义。其多视角对象关联机制结合了观测融合与语义对比学习,以提高身份一致性和语义判别性。边界重建优化进一步细化局部边界结构,以提升轮廓质量。在LERF和3D-OVS上的实验展示了强大的定量和定性性能。我们的方法在LERF上达到了59.1%的mIoU,相比基线获得了13.4%的相对提升。项目页面:https://nora202.github.io/hybridgaussians。
英文摘要:
Open-vocabulary 3D segmentation localizes objects from free-form text queries, but remains challenging in real image sequences: incomplete or noisy 2D supervision destabilizes multi-view identity assignment, while full-scene semantic learning weakens object-level discriminability. We introduce Hybrid Gaussians, a unified 3D representation jointly modeling object association and language-aligned semantics. Its Multi-View Object Association mechanism combines Observation Fusion and Semantic Contrastive Learning to improve identity consistency and semantic discrimination. Boundary Reconstruction Optimization further refines local boundary structure to improve contour quality. Experiments on LERF and 3D-OVS demonstrate strong quantitative and qualitative performance. Our method achieves 59.1\% mIoU on LERF, yielding a 13.4\% relative gain over the baseline. Project page: https://nora202.github.io/hybridgaussians.