arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14445cs.CV

棉花-SF YOLO:学习结构和频率线索用于复杂田间环境下棉花幼蕾的早期检测

Cotton-SF YOLO: Learning Structural and Frequency Cues for Early Cotton Square Detection in Complex Field Environments

Chengjia Zhang, Yu Li, Feiri Ali, Yan Zhang, Xin Chen, Longke He, Daokun Ma, Liting Gao

首次发表
浏览论文内容

中文总结 AI 辅助

针对复杂田间环境下棉花幼蕾早期检测难题,提出Cotton-SF YOLO框架,引入动态蛇形卷积和频域特征调制模块,在自建数据集上训练评估,相比基线模型性能提升,证明结构和频率线索互补可实现最佳检测效果。

中文摘要 AI 辅助

棉花幼蕾是棉花早期生殖生长的重要表型指标,其在复杂田间环境下的自动检测对棉花生长监测和精准栽培管理至关重要。然而,由于棉花幼蕾小、常被遮挡、易模糊、受光照变化影响且与周围棉叶对比度低,相关研究不足。为此,提出基于YOLO26m的Cotton-SF YOLO框架。引入动态蛇形卷积提升对小且不规则幼蕾边界的感知,设计频域特征调制模块增强频域表征。在新构建的带注释田间数据集上训练和评估,该模型的mAP$_{50}$、mAP$_{50:95}$和召回率分别为0.8196、0.4942和0.7939,优于基线YOLO26m。消融实验和可视化表明结构和频率线索互补效果最佳。

英文摘要

Cotton squares are important phenotypic indicators of the early reproductive growth of cotton, and automatic field detection of cotton squares provides an important basis for cotton growth monitoring and precision cultivation management. However, early cotton square detection in complex field environments remains insufficiently explored, as cotton squares are small, frequently occluded, easily blurred, subject to illumination variations, and exhibit low contrast against surrounding cotton leaves. To address these challenges, we propose a task-oriented framework based on YOLO26m, named Cotton-SF YOLO, for cotton square detection under natural field conditions. To improve the perception of small and irregular cotton square boundaries, we introduce Dynamic Snake Convolution into the detector, enabling adaptive extraction of deformable edge features. Furthermore, a frequency-domain feature modulation module is designed by incorporating spectral enhancement into the C2f structure, which recalibrate frequency-domain representations and strengthen discriminative edge and texture cues while reducing interference from complex cotton leaf backgrounds. Trained and evaluated on our newly constructed and annotated field dataset with manually annotated cotton squares, the proposed model achieves mAP$_{50}$, mAP$_{50:95}$, and recall values of 0.8196, 0.4942, and 0.7939, improving over the baseline YOLO26m by 1.25%, 3.45%, and 2.96%, respectively. Ablation experiments and visualization demonstrate that the best performance is achieved with the complementary effects of structural and frequency cues.

发表机构

  • School of Aeronautics and Astronautics, Xichang University(西昌学院航空航天学院)
  • School of Information Technology, Xichang University(西昌学院信息技术学院)
  • College of Information and Electrical Engineering, China Agricultural University(中国农业大学信息与电气工程学院)
  • Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(萨里大学视觉、语音和信号处理中心)

机构由 AI 辅助整理,请以论文原文为准。

↑