Scene-Q:基于选择性VLM推理的置信度感知粗到细三维场景查询
Scene-Q: Confidence-Aware Coarse-to-Fine Querying of 3D Scenes with Selective VLM Reasoning
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Scene-Q提出置信度感知的粗到细查询框架,通过温度缩放和选择性VLM推理,在保持快速路径的同时提升开放词汇三维场景理解与检索性能。
AI中文摘要:
室内移动机器人需要开放词汇的场景理解,将自然语言查询锚定在一致的三维地图中。许多现有系统最终依赖对比图像-文本编码器的余弦相似度检索,这在标签近似同义或出现多个相似实例时效率高但脆弱。我们提出Scene-Q,一个置信度感知的粗到细查询框架,通过温度缩放归一化编码器分数,并仅对低置信度情况选择性调用推理VLM。高置信度查询由快速检索回答,而模糊查询则利用原始多视图图像和实例边界框在小规模top-K候选集上重新排序,以低成本实现上下文感知的消歧。Scene-Q在ScanNet200上改进了开放词汇三维实例分割,并在真实世界重建上改进了自然语言三维实例检索,在空间和关系查询上提升最大,同时保持相当比例的查询在快速路径上。
英文摘要:
Indoor mobile robots require open-vocabulary scene understanding that grounds natural-language queries in a consistent 3D map. Many existing systems ultimately rely on cosine-similarity retrieval with contrastive image--text encoders, which is efficient but brittle when labels are near-synonymous or multiple similar instances appear. We present Scene-Q, a confidence-aware coarse-to-fine querying framework that normalizes encoder scores with temperature scaling and selectively invokes a reasoning VLM only for low-confidence cases. High-confidence queries are answered by fast retrieval, while ambiguous ones are reranked over a small top-K candidate set using the original multi-view images and instance bounding boxes, enabling context-aware disambiguation at low cost. Scene-Q improves open-vocabulary 3D instance segmentation on ScanNet200 and natural-language 3D instance retrieval on real-world reconstructions, with the largest gains on spatial and relational queries while keeping a substantial fraction of queries on the fast path.