AI 中文总结
针对无人机目标检测,本文基于YOLO-World框架,用A2C2f层替换C2f层,引入注意力机制和并行处理结构,提升小目标检测性能,在VisDrone数据集实验中各项指标显著优于原模型,提供了高精度检测方案。
AI 中文摘要
基于无人机的目标检测技术发展迅速,研究趋势已从检测预定义目标扩展到识别特定目标对象,如通过文本提示准确检测感兴趣的对象。本文提出一种基于多模态的高效目标检测模型以提高小目标检测性能。该方法基于YOLO-World框架,用基于注意力的A2C2f层替换YOLOv8主干中的C2f层,能更精确表示局部特征,增强计算精度。在VisDrone数据集上的对比实验表明,该模型性能优于原始模型,各项指标均有显著提升,验证了该方法为无人机图像和视频应用环境中的目标检测提供了有效且高精度的解决方案。
英文摘要
Drone-based object detection technology has advanced rapidly, becoming increasingly sophisticated and efficient. Recently, research trends have expanded beyond the detection of predefined objects toward the identification of specified target objects. For example, desired targets can be specified through textual prompts, enabling accurate detection of objects of interest. To address this demand, this paper proposes an efficient multimodal-based object detection model aimed at improving small object detection performance. The proposed method is built upon the YOLO-World framework and replaces the C2f layers used in the YOLOv8 backbone with attention-based A2C2f layers. This modification enables more precise representation of local features, particularly for small objects or objects with well-defined boundaries. In addition, the incorporation of attention mechanisms and parallel processing structures significantly enhances the model's computational accuracy. Comparative experiments conducted on the VisDrone dataset demonstrate that the proposed model outperforms the original YOLO-World model. Specifically, precision increases from 43.0% to 45.1%, recall from 32.8% to 35.0%, the F1 score from 37.2% to 39.4%, mAP@0.5 from 32.5% to 35.2%, and mAP@0.5-0.95 from 18.5% to 19.9%, confirming a substantial improvement in detection accuracy. These results verify that the proposed approach provides an effective and highly accurate solution for object detection in drone-based image and video application environments.
CommentsPublished in the International Journal of Interactive Mobile Technologies