基于空间掩码与通道融合的深度多模态融合检测
Deep Multimodal Fusion Detection through Spatial Mask and Channel Competition
浏览论文内容
中文总结 AI 辅助
本文针对现有多模态融合检测的过拟合问题,提出Attention-Driven Complementarity Resampling框架,通过语义掩码交换与可学习通道竞争实现跨模态目标检测的性能提升,在多数据集上取得了有竞争力的结果。
中文摘要 AI 辅助
深度多模态目标检测通过挖掘模态特性已展现出良好性能,但现有特征级融合方法主要在两种模态间权衡并将其统一到单一表示空间,这会导致双骨干网络架构中单一模态的统计特性出现过拟合或过度专业化。本文提出Attention-Driven Complementarity Resampling框架以稳健提升跨模态目标检测性能:基于共享通道空间注意力机制,我们在训练阶段引入语义掩码交换,主动混合模态边界,迫使骨干网络学习不依赖固定模态标签的通用特征;随后提出可学习的通道竞争,以通道方向、可学习的方式采样并聚合特征。在多个数据集上的实验表明,该方法有效,在现有最先进方法中取得了有竞争力的结果,源代码在补充材料中提供。
英文摘要
Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusion methods mainly weigh between two modalities and unify them in a unified representation space. This can lead to overfitting or over-specialization of the statistical properties of a single modality within a dual-backbone architecture. This paper proposes an Attention-Driven Complementarity Resampling framework for robust improvement of cross-modality object detection. Based on a shared channel spatial attention mechanism, we first introduce the semantic mask exchange to actively mix the boundaries of the modalities during the training phase, forcing the backbone network to learn generalized features without relying on fixed modal labels. Then we propose a learnable channel competition to sample and aggregate features in a channel-wise and learnable way. Our experiments on multiple datasets demonstrate that the proposed method is effective and yields competitive results among existing state-of-the-art approaches. The source code is provided in the supplementary material.
发表机构
- KTH Royal Institute of Technology(KTH皇家理工学院)
- University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。