HDMamba-YOLO:面向无人机小目标的高效状态空间感知与局部空间重建
HDMamba-YOLO: Efficient State-Space Perception and Local Spatial Reconstruction for UAV Small Object
浏览论文内容
中文总结 AI 辅助
提出HDMamba-YOLO,一种阶段式异构SSM-CNN检测器,通过感知-重建-对齐-交互逻辑解决无人机小目标检测,在VisDrone2019和AI-TOD上取得领先性能。
中文摘要 AI 辅助
无人机图像中的小目标检测面临视觉证据薄弱、边界模糊、目标分布密集以及背景复杂等挑战。因此,有效的检测需要长距离上下文信息以进行目标与背景的区分,同时保留明确的局部二维结构以实现精确定位。这些需求出现在检测流程的不同阶段,无法通过统一的特征处理策略自然解决。我们提出了混合双域Mamba-YOLO(HDMamba-YOLO),一种按感知-重建-对齐-交互逻辑组织的阶段式异构SSM-CNN检测器。基于EfficientVMamba的EVSS在骨干网络中建立长距离上下文感知,而PhasePatchMerging2D提供相位感知的分层过渡。DST-Wrapper和Native C3k2-ASSAF随后在FPN/PAN聚合过程中执行从感知到重建的过渡以及重复的局部二维重建。DySample提供内容自适应的跨尺度重采样,而OS-CVTIA引入宏观-微观交互和任务特定调制以进行定位和分类。在VisDrone2019上,HDMamba-YOLO-B实现了42.737%的mAP50和25.713%的mAP50:95,参数为10.042M,修正后的GFLOPs为29.879。HDMamba-YOLO-Lite实现了41.140%的mAP50和24.741%的mAP50:95,参数为5.344M。在统一的AI-TOD评估协议下,HDMamba-YOLO-B获得了21.621%的AP和47.881%的AP50。受控消融实验进一步支持了状态空间感知、卷积重建、动态对齐和任务交互在无人机小目标检测中的阶段式分配。
英文摘要
Small-object detection in UAV imagery is challenged by weak visual evidence, ambiguous boundaries, dense object distributions, and complex backgrounds. Effective detection therefore requires long-range contextual information for target-background discrimination while preserving explicit local two-dimensional structures for accurate localization. These requirements arise at different stages of the detection pipeline and are not naturally addressed by a uniform feature-processing strategy. We propose Hybrid Dual-domain Mamba-YOLO (HDMamba-YOLO), a stage-wise heterogeneous SSM-CNN detector organized according to a perception-reconstruction-alignment-interaction rationale. EfficientVMamba-based EVSS establishes long-range contextual perception in the backbone, while PhasePatchMerging2D provides phase-aware hierarchical transitions. DST-Wrapper and Native C3k2-ASSAF then perform perception-to-reconstruction transition and repeated local two-dimensional reconstruction during FPN/PAN aggregation. DySample provides content-adaptive cross-scale resampling, while OS-CVTIA introduces macro-micro interaction and task-specific modulation for localization and classification. On VisDrone2019, HDMamba-YOLO-B achieves 42.737% mAP50 and 25.713% mAP50:95 with 10.042M parameters and 29.879 corrected GFLOPs. HDMamba-YOLO-Lite achieves 41.140% mAP50 and 24.741% mAP50:95 with 5.344M parameters. Under the unified AI-TOD evaluation protocol, HDMamba-YOLO-B obtains 21.621% AP and 47.881% AP50. Controlled ablations further support the stage-wise allocation of state-space perception, convolutional reconstruction, dynamic alignment, and task interaction for UAV small-object detection.
发表机构
- Nanjing University of Science and Technology(南京理工大学)
机构由 AI 辅助整理,请以论文原文为准。