发表机构
University of Technology Sydney; Kingswood; Guide Dogs New South Wales/Australian Capital Territory(悉尼科技大学; 金斯伍德; 新南威尔士州/澳大利亚首都领地导盲犬协会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出无需训练的VOIM,可从RGB-D或单目RGB构建开放词汇3D实例地图,在ScanNet++、Replica数据集上的mIoU优于或匹配现有在线系统,建图阶段是性能关键因素。
AI 中文摘要
我们提出了体素基在线实例管理器(VOIM,Voxel-Grounded Online Instance Manager),这是一种无需训练的体素基实例管理器,可仅从RGB-D或单目RGB构建开放词汇3D实例地图,而此前的无需训练系统均未涉及该领域。在线系统通常在首次检测到物体实例时就对其进行分割和标注,此时证据最为薄弱;相反,VOIM会推迟标注和实例决策,直到未修改的现成感知系统提供的软证据在不同视图间的每个体素上积累起来。我们表明,建图阶段而非特定的感知模型决定了结果:在ScanNet++数据集上的四种感知配置(改变区域描述符、检测器标签先验和掩码源)下,该地图的平均交并比(mIoU)超过了最强的在线RGB-D系统OVO-SLAM,提升幅度在4.8至11.7之间。感知并非中性,替换该基线的自身描述符系列会损失4.1的提升幅度,不过该基线拥有略优的2D描述符(在三个场景上为33.7,而VOIM为31.5),仍能实现更弱的地图。在同等条件下,VOIM在ScanNet++上达到44.07 mIoU,而基线为32.37,在所有10个场景和两种聚合方式(合并值33.31对25.97)中均获胜;同一系统无需修改即可完全在单目RGB上运行,在Replica数据集上与基线的合并值匹配(27.80对27.50)。该优势是特定场景的:在Replica的全类别评分下,匹配输入给出了拆分结果,合并值为28.60对27.50,而按场景平均值则为24.59对30.11。房间尺度受标签限制,建筑尺度受漂移限制。标注并非实时运行,受全词汇的逐类检测主导。这些地图可导出占用网格,并能对自由形式的查询解析为物体实例。
英文摘要
We present Voxel-Grounded Online Instance Manager (VOIM), a training-free voxel-grounded instance manager that builds open-vocabulary 3D instance maps from RGB-D or from monocular RGB alone, a regime no prior training-free system addresses. Online systems typically segment object instances and label them at first detection, committing when evidence is weakest. VOIM instead defers label and instance decisions until soft evidence from unmodified, off-the-shelf perception has accumulated per voxel across views. We show that the mapping stage, rather than the particular perception models, carries the result: across four perception configurations on ScanNet++, varying the region descriptor, the detector label prior and the mask source, the map exceeds the strongest online RGB-D system, OVO-SLAM, by between 4.8 and 11.7 mIoU. Perception is not neutral, and substituting that baseline's own descriptor family costs 4.1 of the margin, yet the baseline carries the marginally better 2D descriptor (33.7 vs. 31.5 mIoU over three scenes) and still realizes the weaker map. Under a like-for-like protocol VOIM reaches 44.07 mIoU on ScanNet++ against 32.37, winning all ten scenes and both aggregations (pooled 33.31 vs. 25.97), and the same system runs unchanged to fully monocular RGB, matching that baseline pooled on Replica (27.80 vs. 27.50). The advantage is regime-specific: under Replica's all-classes scoring, matched inputs give a split result, 28.60 vs. 27.50 pooled against 24.59 vs. 30.11 on the per-scene mean. Room scale is label-limited and building scale drift-limited. Labeling does not run in real time, dominated by per-class detection over the full vocabulary. The maps export occupancy grids and resolve free-form queries to object instances.