arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22913cs.CV

在杂乱的田间视频中检测小型传粉者

Small-Pollinator Detection in Cluttered Field Video

  • Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Onur Onal, Chen Chen

AI总结:

研究在杂乱田间视频中检测小型传粉者的挑战,通过BuzzSpot数据集比较YOLO和RF-DETR模型,发现RF-DETR Large在1344像素分辨率下效果最佳,检测器选择和分辨率提升更有效,还找出蜜蜂与食蚜蝇区分瓶颈,为后续研究提供方向。

AI中文摘要:

在田间视频中检测传粉者具有挑战性,因为目标小、视觉上相似,且在模糊和遮挡的杂乱植被背景下观察。我们在实际的单GPU计算预算下对小型传粉者检测进行了系统的实证研究。使用BuzzSpot挑战数据集,我们在不同输入分辨率下比较了YOLO和RF-DETR模型,并评估了切片推理、类门控融合、大小路由集成和事后时间处理。RF-DETR Large在1344像素分辨率下取得了最佳隐藏测试结果,达到0.405 mAP50:95,优于1120像素模型(0.379)和最佳单模型YOLO26m基线(0.366)。最大的收益来自采用RF-DETR并提高其输入分辨率,表明检测器选择和输入分辨率比增加推理时间复杂度更有效;分辨率增益对小物体以及较罕见的大黄蜂和蛾类最强。切片推理融合、大小路由集成和热启动1536像素延续未超过此结果,而事后时间处理未改善泄漏诊断评估。错误分析确定蜜蜂与食蚜蝇的区分是最明显的剩余瓶颈:相邻帧很少为事后校正提供正确分类的食蚜蝇证据。这些发现促使在最终分类决策之前进行学习的特征级时间聚合。

英文摘要:

Detecting pollinators in field video is challenging: targets are small, visually similar, and observed against cluttered vegetation under blur and occlusion. We present a systematic empirical study of small-pollinator detection under a practical single-GPU compute budget. Using the BuzzSpot challenge dataset, we compare YOLO and RF-DETR models across input resolutions and evaluate sliced inference, class-gated fusion, size-routed ensembling, and post-hoc temporal processing. RF-DETR Large at 1344-pixel resolution achieved our best hidden-test result, reaching 0.405 mAP50:95 and outperforming the 1120-pixel model (0.379) and the best single-model YOLO26m baseline (0.366). The strongest gains came from adopting RF-DETR and increasing its input resolution, indicating that detector choice and input resolution were more effective levers than added inference-time complexity; the resolution gain was strongest for small objects and the rarer bumblebee and moth classes. Sliced-inference fusion, size-routed ensembling, and warm-started 1536-pixel continuation did not surpass this result, while post-hoc temporal processing did not improve the leaked diagnostic evaluation. Error analysis identified bee-hoverfly discrimination as the clearest remaining bottleneck: neighboring frames rarely supplied correctly classified hoverfly evidence for post-hoc correction. These findings motivate learned feature-level temporal aggregation before the final classification decision.

补充信息

↑