检测器在杂乱操作中检测的是物体存在而非遮挡
Detector Confidence Signals Presence Rather Than Occlusion in Cluttered Manipulation
- School of Computer Science & Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究开放词汇检测器在物体被遮挡时的表现,通过与几何分割预言机配对审核其对遮挡的反映,发现其置信度与可见性无关,会误判,基于置信度的指标和门控有问题,还将效应与校准错误等联系起来并给出建议。
AI中文摘要:
将一个命名物体遮挡至大约八分之一可见,开放词汇检测器对该物体存在的置信度几乎不变;周围杂物增多时置信度甚至上升。在真实视频中,同一类别的另一个实例里,检测器在99%的遮挡帧中仍报告物体存在。这很关键,因为该置信度常被视为可见性信号用于多种任务。通过与能提供无检测器真实可见性的几何分割预言机配对,审核其是否反映遮挡。结果表明,随着真实可见性降低,置信度几乎不变且与可见性无关,检测器在约十分之九的场景中报告目标存在,对同一类别的干扰物也触发。这种情况在三个检测器、九个物体类别、两种模拟器等多种情况下都存在。由此产生两个后果:基于置信度的指标低估了解决遮挡的价值约十倍,基于置信度的门控恰好在物体被隐藏时触发。尝试的单视图信号都无法标记遮挡,因为遮挡物位于目标位置。我们将此效应与检测器校准错误和物体幻觉联系起来,发布了受控基准,并推荐基于目标的信号用于门控和评估。
英文摘要:
Occlude a named object until about an eighth of it remains visible, and an open-vocabulary detector's confidence that the object is present barely changes; as the clutter around it grows the confidence can even rise. On real video the detector still reports the object present in 99% of occluded frames, on another instance of the same category. This matters because that confidence is widely read as a visibility signal, used to threshold detections, evaluate open-vocabulary detectors, ground language, retrieve instances, and gate active perception. We audit whether it reflects occlusion by pairing every view with a geometry-segmentation oracle that gives detector-free ground-truth visibility. As true visibility falls from every scene to one in eight, the confidence stays nearly constant and uncorrelated with visibility, and the detector reports the target present in about nine of ten scenes, firing on same-category distractors: it signals that the category is present somewhere, not that the specific target is visible. The failure holds across three detectors (Grounding DINO, OWLv2, and Segment Anything Model 3), nine object categories, two simulators with different renderers and object sets, built and natural occlusion, and real video. Two consequences follow: a confidence-based metric understates the value of resolving occlusion by about ten times (8 against 88 points in our active-perception setting), and a confidence-based gate fires exactly when the object is hidden. No single-view signal we tried, including a realizable localization check, flags the occlusion, because the occluders sit where the target is. We connect the effect to detector miscalibration and object hallucination, release the controlled benchmark, and recommend target-grounded signals for gating and evaluation.