选择性棉花铃定位用于机器人采摘:田间条件下深度学习视觉模型的评估
Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions
浏览论文内容
中文总结 AI 辅助
本研究评估了多种深度学习视觉模型,发现YOLOv12-m-seg在棉花铃检测与分割中性能最优,适用于机器人选择性采摘,具备田间部署潜力。
中文摘要 AI 辅助
本研究开发并评估了一种基于深度学习的感知框架,用于选择性机器人采棉。数据集包含使用三台相机在不同自然光照和天气条件下采集的1,008张标注田间图像。评估了YOLOv8至YOLOv13系列的目标检测模型,均使用其默认配置,同时使用YOLOv8-seg、YOLOv11-seg、YOLOv12-seg、Segment Anything Model (SAM)、SAMv2.1、FastSAM以及结合Recognize Anything Model (RAM)的Grounded-SAM评估了分割性能。在检测模型中,GELAN-s在平均精度(mAP)和推理速度之间取得了最有利的平衡,其mAP为86.1%,精确率为81.6%,召回率为76.6%,F1分数为79.0%,平均每张图像推理时间为42.3毫秒。在直接分割模型中,YOLOv12-m-seg在AP@0.5和FPS之间提供了最有利的平衡,实现了83.7%的分割AP@0.5,每张图像推理时间为20.4毫秒。在检测提示的分割方法中,由GELAN-s生成的边界框提示改善了SAM和SAMv2.1对棉花铃的定位,而SAMv2.1 Tiny始终优于FastSAM和结合RAM的Grounded-SAM。在与人工标注分割掩膜的基于面积的评估中,YOLOv12-m-seg达到了0.966的$R^2$值,而GELAN-s + SAMv2.1 Tiny为0.860。使用UR5e机器人操纵器、定制末端执行器和ZED2i立体相机进行的田间实验进一步验证了YOLOv12-m-seg模型在不同置信度水平下用于实时棉花铃检测、分割和选择性采摘的有效性。这些结果表明,YOLOv12-m-seg为机器人棉花收获提供了一种高效的感知模型,并具有很强的田间部署潜力。
英文摘要
This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.
发表机构
- University of Georgia(佐治亚大学)
- Mississippi State University(密西西比州立大学)
机构由 AI 辅助整理,请以论文原文为准。