arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HGSQ:用于实时航拍小目标检测的热力图引导稀疏查询检测器

HGSQ: Heatmap-Guided Sparse Query Detector for Real-Time Aerial Small Object Detection

Yangchen Zeng

arXiv 2609.13306首次发表:更新:

发表机构

Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对实时航拍小目标检测,提出热力图引导稀疏查询检测器HGSQ,通过轻量级热力图预算预测器引导查询选择、局部形状细化和自适应解码器预算,在NWPU VHR-10和VisDrone2019上取得高精度并实现实时推理。

AI 中文摘要

实时航拍小目标检测是一个重要的视觉信号和图像处理问题,要求检测器在保留细粒度定位的同时,避免在大背景区域上进行冗余计算。本文聚焦于这种面向部署的航拍/无人机场景,而非声称提出一种适用于所有目标检测场景的通用检测器。现有的基于Transformer的检测器提供了强大的全局建模能力,但其密集查询初始化和多层解码器仍在背景令牌上耗费大量计算,当小目标仅占据稀疏图像区域时,这种效率较低。为解决此问题,本文提出HGSQ,一种用于实时航拍小目标检测的热力图引导稀疏查询检测器。HGSQ使用轻量级热力图预算预测器(HBP)在单次前向传播中预测前景预算图。预测的热力图随后被三个固定组件使用:热力图引导稀疏查询选择(HSQS),从高置信度前景位置初始化解码器查询;热力图门控轻量蛇形卷积(HGLSConv),仅在热力图激活的小目标区域执行局部形状细化;以及自适应查询-解码器预算分配(AQDB),根据估计的目标密度调整查询预算和解码器深度。与事后热力图生成不同,HGSQ在部署时将热力图视为实时计算预算而非可视化图。在NWPU VHR-10和VisDrone2019上的实验表明,HGSQ在NWPU VHR-10上达到95.10 mAP50,在VisDrone2019上达到54.8 mAP50,同时将GFLOPs降至48.6,并在我们的TensorRT FP16部署协议下于RTX 4070上以96.0 FPS运行。

英文摘要

Real-time aerial small object detection is an important visual signal and image processing problem, requiring a detector to preserve fine-grained localization while avoiding redundant computation on large background regions. This paper focuses on this deployment-oriented aerial/UAV setting rather than claiming a universal detector for all object detection scenarios. Existing Transformer-based detectors provide strong global modeling, but their dense query initialization and multi-layer decoder still spend substantial computation on background tokens, which is inefficient when small objects occupy only sparse image regions. To address this problem, this paper proposes HGSQ, a Heatmap-Guided Sparse Query Detector for real-time aerial small object detection. HGSQ uses a lightweight Heatmap Budget Predictor (HBP) to predict a foreground budget map in a single forward pass. The predicted heatmap is then used by three fixed components: Heatmap-Guided Sparse Query Selection (HSQS), which initializes decoder queries from high-confidence foreground positions; Heatmap-Gated Lite Snake Convolution (HGLSConv), which performs local shape refinement only on heatmap-activated small-object regions; and Adaptive Query-Decoder Budgeting (AQDB), which adjusts the query budget and decoder depth according to the estimated object density. Unlike post-hoc heatmap generation, HGSQ treats the heatmap as a real-time computation budget rather than a visualization map during deployment. Experiments on NWPU VHR-10 and VisDrone2019 show that HGSQ achieves 95.10 mAP50 on NWPU VHR-10 and 54.8 mAP50 on VisDrone2019, while reducing GFLOPs to 48.6 and running at 96.0 FPS on an RTX 4070 under our TensorRT FP16 deployment protocol.

Comments3 figures and 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑