arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34853cs.CVcs.AI

EviSplat:在3D高斯泼溅中保留多视角证据用于开放词汇分割

EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation

  • DGIST(大邱庆北科学技术院)
  • Chubu University(中部大学)
  • HUAWEI(华为)
  • KAIST(韩国科学技术院)
  • Elith Inc.(Elith公司)

机构由 AI 辅助整理,请以论文原文为准。

Sungho Moon, Kota Shimomura, Junwoo Park, Wonhyeok Choi, Seunghun Lee, Takayoshi Yamashita, Sunghoon Im

AI总结:

EviSplat通过保留3D高斯泼溅中每个观测特征作为证据,并在查询时按相关性聚合,解决了查询前整合导致线索抑制的问题,实现了开放词汇分割的最先进性能。

AI中文摘要:

开放词汇3D场景理解能够根据自由形式的文本查询实现物体定位和分割,而无需固定的类别词汇表。许多近期方法基于3D高斯泼溅,并在查询已知之前,将多视角观测(例如来自各个视角的掩码裁剪)整合到语言特征或紧凑的物体描述符中。然而,同一物体的观测在不同视角下会有所变化,且信息量并不相同:有些观测揭示了与特定查询相关的线索,而另一些则提供不完整或误导性的证据。因此,查询前的整合可能会抑制后续查询所依赖的线索。我们提出了EviSplat,它将各个观测特征保留为后续文本查询的证据。EviSplat在表示物体、物体部件或背景区域的类无关3D实例中保留单个观测特征。它还针对每个高斯学习一个分布,描述其观测所支持的视觉外观。给定一个文本查询,EviSplat使用每个实例最相关的观测来对其进行评分。然后,它通过将实例级相关性与局部支持的证据相结合,并根据该高斯被观测的频率和明确程度进行加权,为每个高斯计算一个分数。因此,不同的查询可以从相同的保留证据中利用不同的视觉线索。跨多个数据集和评估协议的实验证明了最先进的性能,支持了将多视角证据保留到查询时并根据查询进行聚合的益处。

英文摘要:

Open-vocabulary 3D scene understanding enables object localization and segmentation from free-form text queries without a fixed category vocabulary. Many recent methods build on 3D Gaussian Splatting and consolidate multi-view observations, such as masked crops from individual views, into language features or compact object descriptors before the query is known. However, observations of the same object vary across viewpoints and are not equally informative: some reveal cues relevant to a particular query, whereas others provide incomplete or misleading evidence. Pre-query consolidation can therefore suppress cues on which a later query depends. We introduce EviSplat, which preserves individual observation features as evidence for later text queries. EviSplat retains individual observation features within class-agnostic 3D instances that represent objects, object parts, or background regions. It also learns, for each Gaussian, a distribution describing which visual appearances its observations support. Given a text query, EviSplat scores each instance using its most relevant observations. It then computes a score for each Gaussian by combining instance-level relevance with locally supported evidence, weighted by how often and how unambiguously that Gaussian was observed. Different queries can thus draw on different visual cues from the same preserved evidence. Experiments across diverse datasets and evaluation protocols demonstrate state-of-the-art performance, supporting the benefit of preserving multi-view evidence until query time and aggregating it according to the query.

补充信息

↑