arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

YOLO12-MambaScan:一种具有高频增强和状态空间建模的高效目标检测器

YOLO12-MambaScan: An Efficient Object Detector with High-Frequency Enhancement and State-Space Modeling

Hao Wang

arXiv 2609.13647首次发表:更新:

AI 中文总结

针对航拍小目标检测中高频信息丢失和全局上下文建模难的问题,提出基于YOLO12的检测器,结合高频增强卷积、坐标注意力和Mamba模块,在VisDrone上达到60.0% mAP@50。

AI 中文摘要

无人机(UAV)技术的快速发展使得航拍图像目标检测在自然资源监测、交通管理和灾害响应中日益重要。在航拍图像中检测小目标仍然困难,因为目标仅占据极少数像素,高频线索容易丢失,并且在杂乱场景中难以建模全局上下文。现有检测器通常保留的边缘、角落和纹理信息不足。我们提出了\ours,一种基于YOLO12架构构建的航拍图像检测器。该模型结合了三路径高频增强卷积模块(TriPathHFConv)、感受野坐标注意力卷积(RFCAConv)以及基于Mamba的全局上下文模块。在VisDrone数据集上,输入分辨率为960*960时,我们的方法达到了60.0%的mAP@50和38.6%的mAP@50:95,展示了在小目标检测中良好的精度-效率权衡。基准测试和数据集协议遵循VisDrone挑战赛的设置。

英文摘要

The rapid development of unmanned aerial vehicle (UAV) technology has made aerial-image object detection increasingly important for natural-resource monitoring, traffic management, and disaster response. Detecting small objects in aerial images remains difficult because objects occupy very few pixels, high-frequency cues are easily lost, and global context is hard to model in cluttered scenes. Existing detectors often retain insufficient edge, corner, and texture information. We propose \ours, an aerial-image detector built on the YOLO12 architecture. The model combines a triple-path high-frequency enhancement convolution module (TriPathHFConv), receptive-field coordinate-attention convolution (RFCAConv), and a Mamba-based global-context module. On VisDrone, at an input resolution of 960*960, ours achieves 60.0% mAP@50 and 38.6%mAP@50:95, demonstrating a favorable accuracy--efficiency trade-off for small-object detection. The benchmark and dataset protocol follow the VisDrone challenge setup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑