arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19592cs.CVcs.RO

选择性棉花铃定位用于机器人采摘:田间条件下深度学习视觉模型的评估

Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估了多种深度学习视觉模型,发现YOLOv12-m-seg在棉花铃检测与分割中性能最优,适用于机器人选择性采摘,具备田间部署潜力。

中文摘要 AI 辅助

本研究开发并评估了一种基于深度学习的感知框架,用于选择性机器人采棉。数据集包含使用三台相机在不同自然光照和天气条件下采集的1,008张标注田间图像。评估了YOLOv8至YOLOv13系列的目标检测模型,均使用其默认配置,同时使用YOLOv8-seg、YOLOv11-seg、YOLOv12-seg、Segment Anything Model (SAM)、SAMv2.1、FastSAM以及结合Recognize Anything Model (RAM)的Grounded-SAM评估了分割性能。在检测模型中,GELAN-s在平均精度(mAP)和推理速度之间取得了最有利的平衡,其mAP为86.1%,精确率为81.6%,召回率为76.6%,F1分数为79.0%,平均每张图像推理时间为42.3毫秒。在直接分割模型中,YOLOv12-m-seg在AP@0.5和FPS之间提供了最有利的平衡,实现了83.7%的分割AP@0.5,每张图像推理时间为20.4毫秒。在检测提示的分割方法中,由GELAN-s生成的边界框提示改善了SAM和SAMv2.1对棉花铃的定位,而SAMv2.1 Tiny始终优于FastSAM和结合RAM的Grounded-SAM。在与人工标注分割掩膜的基于面积的评估中,YOLOv12-m-seg达到了0.966的$R^2$值,而GELAN-s + SAMv2.1 Tiny为0.860。使用UR5e机器人操纵器、定制末端执行器和ZED2i立体相机进行的田间实验进一步验证了YOLOv12-m-seg模型在不同置信度水平下用于实时棉花铃检测、分割和选择性采摘的有效性。这些结果表明,YOLOv12-m-seg为机器人棉花收获提供了一种高效的感知模型,并具有很强的田间部署潜力。

英文摘要

This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.

发表机构

  • University of Georgia(佐治亚大学)
  • Mississippi State University(密西西比州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑