arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越对称融合:利用任务相关的模态优势进行RGB-事件小目标检测

Beyond Symmetric Fusion: Exploiting Task-Dependent Modality Strengths for RGB-Event Small Object Detection

Ziheng Wang, Chaolang Li, Yutong Yang, Xiaohan Xu, Chongxiang Yang, Hengxuan Zhong, Zhen Liang, Pengwen Dai

arXiv 2608.01302首次发表:更新:

AI 中文总结

该研究针对现有RGB-事件检测器的对称融合设计缺陷,提出AERODet模型,通过SURE和TDSR实现任务相关的模态利用,在FRED和NeRDD数据集上取得最优性能,FRED挑战性拆分提升10.7 mAP点。

AI 中文摘要

当前最先进的RGB-事件检测器通过融合RGB和事件数据的互补特征提升了对小型、快速移动物体的检测性能,但它们通常将两种模态融合为统一表示用于定位和分类,这种任务对称设计与两种模态应根据其任务特定优势发挥不同作用的直觉不符。为研究该问题,我们开展了模态特定评估,发现两种模态的相对优势会随任务反转:事件数据对类别无关定位更有效,而RGB数据在定位目标区域内提供更强的类别证据。受此任务依赖不对称性启发,我们提出了不对称事件-RGB目标检测Transformer(AERODet)。在类别无关定位阶段,尺度感知不确定性可靠性估计(SURE)从两种模态的目标性响应热图计算其相对可靠性,并在解码器聚合多模态特征时相应校准它们的贡献;获得候选框后,任务解耦语义细化(TDSR)将分类与定位解耦,使用RGB感兴趣区域(RoI)特征进行细粒度分类。在FRED和NeRDD数据集上的大量实验表明,AERODet达到了最先进的性能,尤其在FRED挑战性拆分上,它比最强的RGB-事件基准高出10.7个平均精度(mAP)点。

英文摘要

State-of-the-art RGB-Event detectors improve the detection of small, fast-moving objects by combining complementary features from RGB and Event data, yet they typically fuse the two modalities into a unified representation for both localization and classification. Such a task-symmetric design is inconsistent with the intuition that the two modalities should play different roles according to their task-specific strengths. To examine this issue, we conduct a modality-specific evaluation and find that the relative advantage of the two modalities reverses across tasks: Event data are substantially more effective for class-agnostic localization, whereas RGB data provide stronger category evidence within localized target regions. Motivated by this task-dependent asymmetry, we propose an Asymmetric Event-RGB Object Detection Transformer (AERODet). During class-agnostic localization, Scale-wise Uncertainty-aware Reliability Estimation (SURE) calculates the relative reliability of the two modalities from their objectness response heatmaps and accordingly calibrates their contributions when the decoder aggregates multimodal features. Once the candidate boxes are obtained, Task-Decoupled Semantic Refinement (TDSR) decouples classification from localization and uses RGB RoI features for fine-grained classification. Extensive experiments on FRED and NeRDD demonstrate that AERODet achieves state-of-the-art performance. In particular, it surpasses the strongest RGB-Event baseline by 10.7 mAP points on the FRED challenging split.

Comments3 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑