arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开放词汇检测中的置信度分数是尺度和语义的有偏混合

Confidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and Semantics

Yi Tang Soon, Jun-Wei Hsieh

arXiv 2607.10993首次发表:更新:

发表机构

Institute of Intelligence Systems, National Yang Ming Chiao Tung University; College of Artificial Intelligence and Green Energy, National Yang Ming Chiao Tung University(国立阳明交通大学智能系统研究所; 国立阳明交通大学人工智能与绿色能源学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究开放词汇检测中置信度分数问题,通过实验表明其是尺度和语义的有偏混合,两种偏差源于CLIP图像级预训练,阈值调整无法消除,无参数温度缩放校正可部分改善小物体召回率并揭示相关限制。

AI 中文摘要

诸如CLIP这样的基础模型使得通过视觉-语言相似性推广到新类别的开放词汇目标检测器成为可能。然而,这些检测器产生的置信度分数并非可靠的定位概率估计:它们将视觉尺度和语义查询特异性与真实检测信号混为一谈。通过对基于三个基础模型的检测器(GroundingDINO、OWL-ViT、YOLO-World)在COCO上进行控制实验,并使用GroundingDINO在LVIS(1203个类别)上进一步复制尺度偏差发现,我们表明s=cos(v,t)是两种效应的有偏混合。尺度偏差(α = +0.064,r = 0.579,p = 1.29 x 10^-58)系统性地提高大物体的分数。语义偏差(β = -0.705,p = 5.23 x 10^-41)抑制通用查询的分数。两种偏差在结构上都不可避免地源于CLIP的图像级预训练。阈值调整无法消除它们:针对每个尺度的最优阈值调整对小物体产生的Delta F1为+0.001,而对大物体为+0.102。无参数温度缩放校正将小物体的Recall@10提高了19.6%(p < 0.01)且无需重新训练。这在池化排序精度上有适度的、可测量的代价,因此偏差在推理时只能部分而非完全可逆。这些发现揭示了将图像级基础模型应用于区域级检测任务的一个基本限制。

英文摘要

Foundation models such as CLIP have enabled open-vocabulary object detectors that generalise to novel categories via vision-language similarity. However, the confidence scores these detectors produce are not reliable localization probability estimates: they conflate visual scale and semantic query specificity with the true detection signal. Through controlled experiments on COCO across three foundation-model-based detectors (GroundingDINO, OWL-ViT, YOLO-World), with the scale-bias finding further replicated on LVIS (1,203 categories) using GroundingDINO, we show that s=cos(v,t) is a biased mixture of two effects. Scale bias (alpha = +0.064, r = 0.579, p = 1.29 x 10^-58) systematically inflates scores for large objects. Semantic bias (beta = -0.705, p = 5.23 x 10^-41) suppresses scores for generic queries. Both biases are structurally inevitable from CLIP's image-level pretraining. Threshold adjustment cannot remove them: oracle per-scale thresholding yields Delta F1 = +0.001 for small objects versus +0.102 for large. A parameter-free temperature scaling correction improves small-object Recall@10 by 19.6% (p < 0.01) without retraining. This comes at a modest, measurable cost to pooled-ranking precision, so the bias is partially, not freely, reversible at inference time. These findings reveal a fundamental limitation of adapting image-level foundation models to region-level detection tasks.

CommentsICPR Workshop 2026 (FMVA)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑