arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FSDC-DETR:面向小目标检测的频域-空域协同DETR

FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection

Aiwen Liu, Chengguang Zhu, Gang Wang, Dandan Zhu, Haodong Lin, Yan Wang, Huiyu Zhou, Zhengyi Pan

arXiv 2607.05176首次发表:更新:

发表机构

Micro-Intelligence; East China Normal University; University of Leicester; Chongqing Normal University(微智感知; 华东师范大学; 莱斯特大学; 重庆师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有检测器易丢失小目标高频分量的问题,提出频域-空域协同DETR框架,通过双分支融合、交互机制与动态下采样实现小目标检测性能显著提升。

AI 中文摘要

小目标检测(SOD)在实际应用中仍是一项极具挑战性的任务。尽管近年相关研究取得了一定进展,现有检测器仍受限于刚性处理机制:该机制将空域聚合与隐式频率混叠、截断过程相互纠缠,导致小目标检测所需的高频分量无法得到充分保留。为解决这些局限性,本文提出频域-空域协同检测Transformer(FSDC-DETR),这是一种可显式建模互补空域与频域表征的全新协同框架。具体而言,我们首先引入双分支频域-空域自适应融合(DBFSAF)模块,以提升频率多样性并自适应捕获频域-空域的判别性表征。基于上述表征,本文进一步在混合编码器中探索频域-空域交互方案,实现向解码器的渐进式特征传播。特别地,通过分流频域-空域特征融合(SFS-FF)实现结构感知的频域-空域聚合,在频域表征与空域表征之间建立双向交互与渐进式跨尺度传播,以完成连贯的判别性建模。同时,通过频域-空域动态下采样(FSD-Down)在尺度转换过程中保留信息丰富的高频响应,最大程度降低多尺度融合过程中的频率退化问题,实现精准的小目标检测。实验结果表明,FSDC-DETR取得了当前最优性能,在VisDrone-DET2019数据集上提升6.4个点的AP,在AITODv2数据集上提升6.6个点的AP,对应小目标的AP增益分别达到6.8和6.9。相关代码可通过本HTTP链接获取。

英文摘要

Small object detection (SOD) remains a challenging task in real-world applications. Despite recent advances, existing detectors remain limited by rigid processing that entangle spatial aggregation with implicit frequency aliasing and truncation, leading to inadequate preservation of high-frequency components for SOD. To tackle these limitations, we propose a Frequency-Spatial Domain Collaborative Detection Transformer (FSDC-DETR), a novel collaborative framework that explicitly models complementary spatial and frequency representations. Specifically, we first introduce Dual-Branch Frequency-Spatial Adaptive Fusion (DBFSAF) to enhance frequency diversity and adaptively capture frequency-spatial domain discriminative representations. Building on these representations, a frequency-spatial interaction scheme is further explored within the hybrid encoder to enable progressive feature propagation to the decoder. In particular, structure-aware frequency-spatial aggregation is achieved through Shunt Frequency-Spatial Feature Fusion (SFS-FF), establishing bidirectional interaction and progressive cross-scale propagation between frequency and spatial representations for coherent discriminative modeling. Meanwhile, informative high-frequency responses are preserved during scale transitions through Frequency-Spatial Dynamic Downsampling (FSD-Down), thereby minimizing frequency degradation throughout multi-scale fusion for the precise SOD. Experimental results demonstrate that FSDC-DETR achieves state-of-the-art performance, improving AP by 6.4 on VisDrone-DET2019 and 6.6 on AITODv2, with gains of 6.8 and 6.9 AP for small objects. The code is available at github.com/nevereverinsomnia/FSDC-DETR.

CommentsAccepted by ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑