arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OV3D-Bench:开放词汇单目3D检测的诊断基准

OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection

Mariia Gladkova, Neehar Peri, Ishan Khatri, Deva Ramanan, Daniel Cremers

arXiv 2608.17110首次发表:更新:

发表机构

Technical University of Munich; Carnegie Mellon University; StackAV(慕尼黑工业大学; 卡内基梅隆大学; StackAV)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出OV3D-Bench诊断基准,评估发现开放词汇单目3D检测器存在语义误标、对提示敏感等问题,且目标感知协议掩盖误差,同时证明用SigLIPv2重映射闭词汇检测器预测可与专门方法竞争,指出语义是主要瓶颈。

AI 中文摘要

开放词汇单目3D检测器在领域内表现出色,但各采用不同评估协议,部分依赖部署时不可用的单图像类别先验,且均将几何与语义信息合并为单一AP指标。为解决此问题,本文提出OV3D-Bench,这是一个诊断基准,在贴近部署的条件下,针对7个室内外数据集对开放词汇单目3D检测器进行比较。该基准将单图像类别名称先验替换为测试时的数据集级别类别名称提示,并沿定位、语义鲁棒性和跨域迁移三个维度解耦检测精度。本文评估了7个代表性检测器,发现:(i)它们能较好地定位对象,但常将正确定位的框误标为语义相邻类别;(ii)精度对提示措辞高度敏感,例如WildDet3D在使用“汽车的详细高分辨率照片”而非“汽车”作为提示时,性能从18.6 AP降至5.4;(iii)广泛采用的目标感知协议会掩盖这些误差,例如在ScanNet上使DetAny3D的AP膨胀1.9倍。最后,本文证明,使用SigLIPv2等对比视觉-语言编码器重新映射冻结闭词汇检测器的预测,可与近期专门构建的开放词汇方法竞争,这表明几何定位已较为成熟,而开放词汇语义仍是主要瓶颈。

英文摘要

Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image category oracles unavailable at deployment, and all collapse geometry and semantics into a single AP metric. To address this, we introduce OV3D-Bench, a diagnostic benchmark that compares open-vocabulary monocular 3D detectors under deployment-realistic conditions across seven indoor and outdoor datasets. Our benchmark replaces the per-image class name oracle with test-time dataset-level class name prompts, and decouples detection accuracy along three axes: localization, semantic robustness, and cross-domain transfer. We evaluate seven representative detectors and find that (i) they localize objects well yet often mislabel a correctly localized box as a semantically adjacent category; (ii) accuracy is highly sensitive to prompt phrasing (e.g. WildDet3D's performance collapses from 18.6 to 5.4 AP when prompted with "a detailed high-resolution photo of a car" rather than "car"); and (iii) the widely adopted target-aware protocol hides these errors (e.g. inflating DetAny3D's AP by 1.9 $\times$ on ScanNet). Lastly, we demonstrate that simply remapping a frozen closed-vocabulary detector's predictions using a contrastive vision-language encoder such as SigLIPv2 performs competitively against recent purpose-built open-vocabulary methods. This indicates that geometric localization is more mature, while open-vocabulary semantics remains the primary bottleneck.

CommentsAccepted to OpenSUN3D workshop at ECCV'26; benchmark is released on https://github.com/mgladkova/ov3d-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑