arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07802cs.CV

面向野外西方蓝鸲检测的基准测试

Towards benchmarking Western Bluebird detection in the wild

Estela Monserrat Arriaga Santana, Julian Rosas Scull, Ibeth P. Alarcón, Bibiana Montoya, Aylin Sosa Mejía, Hugo Jair Escalante

首次发表
浏览论文内容

中文总结 AI 辅助

针对野外西方蓝鸲检测,构建含6000余张4K图像的新基准,评估监督与开放词汇模型,发现监督方法最优,微调可提升开放词汇性能,强调领域适应在小型物体检测中的关键作用。

中文摘要 AI 辅助

在自然环境中进行鸟类监测具有挑战性,因为某些鸟类相对于场景而言体型较小、背景杂乱、光照变化以及观察者视角不同。由于缺乏大规模、真实的数据集(这些数据集对于理解行为模式至关重要),进展进一步受限。为弥补这一空白,我们引入了一个新的基准数据集,用于西方蓝鸲(Sialia Mexicana)的检测和分割,该数据集包含来自41个录制会话的超过6000张标注图像。该数据集具有高分辨率(4K)野外图像,其中鸟类仅占据图像的一小部分。我们评估了监督式检测器、在零样本和微调设置下的开放词汇模型,以及分割方法。监督式检测器总体上仍然最可靠,其中Faster R-CNN实现了最高的检测mAP,而RT-DETR提供了最佳的精确率-召回率权衡。开放词汇模型在零样本设置下表现不佳;然而,微调显著提升了它们的性能,YOLO-World变得与监督方法具有竞争力,并实现了最高的精确率、F1分数和mAP@0.5。在分割方面,监督方法显著优于Grounded-SAM和SAM 3:Mask R-CNN实现了最高的掩码mAP,而YOLOv8-Seg提供了最佳的精确率和最快的推理速度。诊断分析进一步表明,失败并非仅由物体大小解释,而是由表观尺度、亮度、对比度、杂乱度、模糊度、拥挤度以及记录会话变化共同作用所致。总体而言,我们的研究结果突显了在杂乱生态场景中进行零样本鸟类检测的难度,并强调了在小型物体环境中领域适应的重要性。

英文摘要

Bird monitoring in natural environments is challenging due to the small size of some species of birds relative to the scene, background clutter, variability in illumination, and the observers' viewpoint. Progress is further limited by the scarcity of large-scale, realistic datasets, which are essential for understanding behavioral patterns. To address this gap, we introduce a new benchmark dataset for the detection and segmentation of Western bluebirds (Sialia Mexicana), comprising over 6,000 labeled images from 41 recording sessions. The dataset features high-resolution (4K) in-the-wild images in which birds occupy only a small fraction of the image. We evaluated supervised detectors, open-vocabulary models under zero-shot and fine-tuned settings, and segmentation approaches. Supervised detectors remain the most reliable overall, with Faster R-CNN achieving the highest detection mAP and RT-DETR offering the best precision-recall trade-off. Open-vocabulary models perform poorly in zero-shot settings; however, fine-tuning substantially improves their performance, with YOLO-World becoming competitive with supervised methods and achieving the highest precision, F1-score, and mAP@0.5. For segmentation, supervised methods significantly outperform Grounded-SAM and SAM 3: Mask R-CNN achieves the highest mask mAP, while YOLOv8-Seg provides the best precision and fastest inference. A diagnostic analysis further shows that failures are not explained by object size alone, but by a combination of apparent scale, brightness, contrast, clutter, blur, crowding, and recording-session variation. Overall, our findings highlight the difficulty of zero-shot bird detection in cluttered ecological scenes and underscore the importance of domain adaptation in small-object settings.

发表机构

  • Universidad Nacional Autónoma de México(墨西哥国立自治大学)
  • Universidad Autónoma de Tlaxcala(特拉克斯卡拉自治大学)
  • The University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
  • INAOE(国家天体物理、光学与电子学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑