发表机构
Tianjin University; Shanghai Jiaotong University; Shanghai AI Laboratory; Shenzhen University of Advanced Technology(天津大学; 上海交通大学; 上海人工智能实验室; 深圳理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对航空目标检测中尺度与密度变化问题,提出LiG-DETR框架,通过共享编码器获取局部放大特征并重组,结合选择性重组与自适应查询分配,在保持大目标性能的同时显著提升小目标检测精度。
AI 中文摘要
航空目标检测面临显著的尺度与密度变化。小目标容易因下采样和特征压缩而退化,而中大型目标则需要充分的全局上下文。现有方法主要遵循两种范式:图像切片提供更清晰的局部证据,但依赖于独立的裁剪级预测和后处理;而特征级和查询级优化保持了统一推理,但操作于已压缩的全图表示,限制了细粒度信息的恢复。这引出一个关键问题:航空检测能否在特征退化之前直接获取高保真的局部证据,并将其整合到一个统一的端到端框架中?为此,我们提出了LiG-DETR,一种高效的全局-局部重组框架,将图像切片重新定义为高保真局部特征获取。一个共享编码器提取全局和局部放大特征,并将其投影到检测器特征空间中。投影后的局部特征根据其原始空间位置进行重组,形成全局对齐的局部特征层,并由单个DETR解码器联合解码全局和局部特征。为减少冗余计算,上下文保留的选择性重组将高分辨率编码聚焦于信息丰富区域,同时保持密集特征布局,而密度感知的自适应查询分配利用编码器提议分数调整解码器查询预算。实验表明,在小目标和中等目标上取得了显著提升,同时保持了强的大目标性能,具有有利的精度-效率权衡和改善的跨域泛化能力。代码将发布。
英文摘要
Aerial object detection faces substantial scale and density variations. Small objects are easily degraded by downsampling and feature compression, while medium and large objects require sufficient global context. Existing methods mainly follow two paradigms: image slicing provides clearer local evidence but relies on independent crop-level prediction and post-processing, whereas feature- and query-level optimization preserves unified inference but operates on already compressed full-image representations, limiting recovery of fine-grained information. This raises a key question: can aerial detection directly acquire high-fidelity local evidence before feature degradation and integrate it into a unified end-to-end framework? To this end, we propose LiG-DETR, an Efficient Global-Local Reassembly framework that reformulates image slicing as high-fidelity local feature acquisition. A shared encoder extracts global and locally magnified features, which are projected into the detector feature space. The projected local features are reassembled according to their original spatial locations to form a globally aligned local feature level, and a single DETR decoder jointly decodes global and local features. To reduce redundant computation, Context-Preserved Selective Reassembly focuses high-resolution encoding on informative regions while preserving a dense feature layout, and Density-Aware Adaptive Query Allocation adapts the decoder query budget using encoder proposal scores. Experiments show substantial gains on small and medium objects while retaining strong large-object performance, with favorable accuracy--efficiency trade-offs and improved cross-domain generalization. The code will be released.