arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21822cs.CV

目标检测基准不完整:标签错误与标注不确定性的作用

Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

Sarina Penquitt, Jonathan Klees, Antonia van Betteray, Parssa Jashnieh, Peter Stehr, Matthias Rottmann, Lars Schmarje

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示目标检测基准因标注缺失而不完整,提出高召回率、软标签聚合的标注流程,并构建不确定性感知与标签错误检测基准,强调未来评估需考虑不确定性。

中文摘要 AI 辅助

尽管目标检测通过改进的架构和开放词汇模型取得了进展,但我们提供了强有力的证据表明,基准质量受到标注不完整性的限制。在四个广泛使用的数据集(COCO、Pascal VOC、Cityscapes、KITTI)上,重新标注揭示了标注对象数量的显著增加(例如,KITTI上最多增加+60%,COCO上增加+40%),这主要是由先前未标注的小型、遮挡或密集排列的实例驱动的。虽然一些差异源于数据集特定的标注约定,但我们一致发现,缺失标注是所有数据集中标签错误的主要来源。为了实现高数据质量,我们引入了一个可扩展的标注流程,该流程强调高召回率,并通过每个对象至少11名标注者的软标签聚合来捕捉不确定性。由此产生的标注提高了覆盖率,并与人类校准良好对齐。我们表明,基准性能对标注质量高度敏感,尽管模型排名基本保持稳定。我们引入了两个大规模基准:(i)一个不确定性感知的目标检测基准,以及(ii)一个基于真实标签错误的标签错误检测基准。我们表明,当前检测器强烈依赖于标注质量,并与人类感知不一致。当前的标签错误检测方法,虽然在合成噪声上表现良好,但在真实标签错误上难以实现高召回率和精确率。我们的结果强调,未来的目标检测基准需要超越确定性标注,转向高召回率、不确定性感知的评估,以最大化有效实例并更好地反映现实世界中的模糊性。

英文摘要

While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.

发表机构

  • Osnabrück University(奥斯纳布吕克大学)
  • University of Wuppertal(伍珀塔尔大学)
  • Wayve

机构由 AI 辅助整理,请以论文原文为准。

↑