arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SuppreSensing:面向多模态目标检测的专家引导特征重校准与差异增强

SuppreSensing: Expert-Guided Feature Recalibration and Discrepancy Augmentation for Multimodal Object Detection

Xin Wu, Zhenyu Gao, Qiankun Zhang, Shaoyong Guo

arXiv 2608.20944首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对遥感多模态目标检测的语义异质性与模态噪声问题,本文提出SuppreSensing,通过EMFR模块、差异增强策略和ECFP模块实现最优性能,在多个数据集上达到SOTA并具备良好泛化性。

AI 中文摘要

遥感领域的多模态目标检测面临语义异质性和模态特定噪声干扰的挑战。为此,本文提出SuppreSensing,将多模态融合重新表述为选择性协作过程,联合建模共享信息与模态特定线索。SuppreSensing首先设计专家驱动的多模态特征重校准(EMFR)模块,将共享共识提取重新表述为输入自适应的多专家选择过程,以缓解多模态融合中的对称陷阱。作为补充,采用模态特定属性增强策略,通过建模双向差异模式增强特定模态特征,减轻跨模态异质性。此外,本文提出基于“专项检查-综合分析-诊断更新”体检范式的专家驱动定制特征纯化(ECFP)模块,迭代过滤冗余信息并强化任务相关语义。在DroneVehicle和VEDAI数据集上的大量实验表明,SuppreSensing达到了最先进的检测性能;在自然场景数据集(FLIR和LLVIP)上的跨域评估进一步验证了其在不同环境条件下的卓越鲁棒性和泛化能力。

英文摘要

Multimodal object detection in remote sensing faces challenges due to semantic heterogeneity and modality-specific noise interference. To this end, we propose SuppreSensing, which reformulates multimodal fusion as a selective collaboration process that jointly models shared information and modality-specific cues. SuppreSensing first designs an Expert-driven Multimodal Feature Recalibration (EMFR) module, which reformulates shared-consensus extraction as an input-adaptive multi-expert selection process to alleviate the symmetry trap in multimodal fusion. Complementing this, a modality-specific attribute augmentation strategy is employed to enhance specific modality features by modeling bidirectional discrepancy patterns, mitigating cross-modal heterogeneity. Furthermore, we propose an Expert-driven Customized Feature Purification (ECFP) module based on a "specialized inspection-comprehensive analysis-diagnostic update" physical examination paradigm to iteratively filter redundancies and reinforce task-relevant semantics. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate that SuppreSensing achieves state-of-the-art detection performance. Cross-domain evaluations on natural scene datasets (FLIR and LLVIP) further validate its superior robustness and generalization capability across diverse environmental conditions.

Comments10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑