arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16309cs.CV

SaaF:面向交互式现实世界对象检索的特定场景模糊感知3D语言场

SaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object Retrieval

Yuga Yano, Daiju Kanaoka, Hakaru Tamukoh, Yasutomo Kawanishi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现实世界场景中交互式对象检索问题,提出SaaF方法。通过引入度量学习策略构建统一特征空间,增强实例辨别性与模糊感知能力,提升了检索准确性,还能稳健处理查询模糊性。

中文摘要 AI 辅助

我们提出了特定场景模糊感知3D语言场(SaaF),这是一种基于高斯点云的新型3D语言场,用于在给定的现实世界场景中进行交互式对象检索。使用自然语言进行交互式对象检索是服务机器人在复杂现实世界环境中运行的关键能力。虽然最近用于对象检索的3D语言场方法在渲染像素和自动编码器压缩的CLIP特征之间建立了关联,但存在两个局限性:由于特征压缩导致相似对象之间的可辨别性降低,以及对模糊查询处理不佳。为了解决这些局限性,SaaF引入了一种度量学习策略来构建一个统一的特征空间,该空间既具有实例辨别性又具有模糊感知能力。为了增强实例级视觉辨别能力,SaaF采用度量学习将同一对象的多个视点的图像特征在特征空间中拉近。为了建立模糊感知,模型对由所提出的方法从每个跟踪对象图像序列生成的多个文本标签进行联合训练,包括模糊描述,以学习目标场景中模糊和特定特征之间的语义关系。该特征空间能够实现细粒度视觉理解,同时允许系统估计查询模糊性并在需要时交互式请求澄清。实验结果表明,SaaF不仅比以前的方法提高了检索准确性,而且在开放词汇设置下能够稳健地检测和处理用户文本查询中的模糊性。

英文摘要

We propose Scene-specific Ambiguity-aware 3D Language Fields (SaaF), a novel Gaussian Splatting-based 3D language field designed for interactive object retrieval in a given real-world scene. Interactive object retrieval using natural language is a crucial capability for service robots operating in complex real-world environments. While recent 3D language field methods for object retrieval establish associations between rendered pixels and autoencoder-compressed CLIP features, they suffer from two limitations: (1) reduced discriminability among similar objects due to feature compression, and (2) poor handling of ambiguous queries, often resulting in unstable or incorrect retrieval. To address these limitations, SaaF introduces a metric learning strategy to construct a unified feature space that is both instance-discriminative and ambiguity-aware. (i) To enhance instance-level visual discrimination, SaaF employs metric learning that pulls image features from multiple viewpoints of the same object closer together in the feature space. (ii) To establish ambiguity awareness, the model jointly trains on multiple text labels generated by the proposed method from each tracked object image sequence, including ambiguous descriptions, to learn the semantic relationships between ambiguous and specific features in a target scene. This feature space enables fine-grained visual understanding while allowing the system to estimate query ambiguity and interactively request clarification when needed. Experimental results demonstrate that SaaF not only improves retrieval accuracy over previous methods but also robustly detects and handles ambiguity in the user text queries under open-vocabulary settings.

发表机构

  • Kyushu Institute of Technology(九州工业大学)
  • RIKEN(理化学研究所)
  • Research Center for Neuromorphic AI Hardware(神经形态人工智能硬件研究中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑