AI 中文总结
本文提出OVIP-SG框架,通过VLM等技术实现小型细粒度物体的实例保留映射与语言引导检索,在Replica数据集上的多项指标优于现有方法,且经实际机器人实验验证有效。
AI 中文摘要
将开放词汇感知融入物体级三维场景图是一把双刃剑:视觉语言检测器虽能恢复长尾类别及闭集模型忽略的小型细粒度物体,但易分割大表面、将小物体合并至相邻大物体,损害实例级一致性并降低映射保真度;现有方法还难以检索未映射目标或判断查询物体是否缺失,阻碍鲁棒具身开放世界导航与探索。本文提出OVIP-SG,这是一个用于实例保留语义映射、功能场景划分及语言引导小型细粒度物体检索的统一框架。OVIP-SG采用视觉语言模型(VLM)枚举场景特定类别以实现鲁棒开放世界检测,通过对称三维交并比(IoU)关联与面积加权特征融合保留小型独立实例,同时利用VLM推断的物体功能将场景划分为紧凑的功能搜索区域;其四阶段级联检索管道进一步结合体素投票,并基于探索覆盖范围判断目标是否缺失。在Replica数据集的统一评估协议下,OVIP-SG的类别平均精度(mAcc)较ConceptGraphs提升6.31个百分点、频率加权平均交并比(F-mIoU)提升5.15个百分点,且类别无关原生实例全景质量(PQ)达0.398;它将搜索区域缩小至室内地面空间的21.8%,物体存在分类的平衡精度达0.773,实际机器人实验进一步验证了其实用有效性。
英文摘要
Integrating open-vocabulary perception into object-level 3D scene graphs is a double-edged sword. While vision-language detectors recover long-tail categories and small, fine-grained objects overlooked by closed-set models, they also tend to fragment large surfaces and merge small objects into larger neighboring objects, compromising instance-level consistency and undermining mapping fidelity. Moreover, existing methods struggle to retrieve previously unmapped targets or determine whether a queried object is absent, hindering robust embodied open-world navigation and exploration. We present OVIP-SG, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval. OVIP-SG uses a vision-language model (VLM) to enumerate scene-specific categories for robust open-world detection. Symmetric 3D Intersection over Union (IoU) association and area-weighted feature fusion preserve small independent instances, while VLM-inferred object functions partition scenes into compact functional search regions. A four-stage cascaded retrieval pipeline further incorporates voxel voting and determines target absence from exploration coverage. Under a unified evaluation protocol on Replica, OVIP-SG outperforms ConceptGraphs by 6.31 points in class-mean accuracy (mAcc) and 5.15 points in frequency-weighted mIoU (F-mIoU) while achieving a class-agnostic native-instance Panoptic Quality (PQ) of 0.398. It reduces the search area to 21.8% of the indoor floor space and reaches 0.773 balanced accuracy for object-presence classification. Real-world robotic experiments further demonstrate its practical effectiveness. Code is available at https://github.com/Agibot-Spatial-Intelligence/OVIP-SG.
Comments15 pages, 6 figures, including appendix