发表机构
Xidian University; Nanyang Technological University; Zhengzhou University; China University of Geosciences; Swinburne University of Technology(西安电子科技大学; 南洋理工大学; 郑州大学; 中国地质大学; 斯威本科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对红外-可见光目标检测中模态缺失及异质模态语义相关性估计难题,提出FlexibleFusion方法,结合MAEC机制与RSPEOT,实现跨模态自适应融合,在任意模态配置下均表现稳定。
AI 中文摘要
红外-可见光目标检测(IVOD)整合可见光与红外传感器的互补证据,用于复杂场景下的可靠感知。实际应用中,传感器可能出现故障或丢帧,导致某一模态不可用或间歇性存在。现有IVOD方法假设两种模态始终存在,固定融合方式在某一模态缺失时会失效。此外,跨异质模态可靠估计语义相关性仍是关键挑战,尤其在光谱分布差异显著的情况下。本文提出FlexibleFusion,一种统一的自适应方法,可灵活分配融合路径与强度,在模态完整和缺失两种场景下无缝运行。其核心是模态感知专家协作(MAEC)机制,选择性激活并聚合跨模态或模态内专家路径,在模态完整时允许跨模态融合,在模态缺失时回退至自融合。此外,本文设计残差自步调熵最优传输(RSPEOT),从传输视角对齐异质特征分布。与标准熵最优传输(EOT)依赖固定稀疏系数不同,RSPEOT引入残差驱动的自步调更新,优先选择可靠匹配并逐步优化较难匹配,该设计在保留可靠语义对齐的同时,减轻了标准EOT的额外优化负担。在模态完整和缺失协议下的综合实验表明,该方法在任意模态配置下均表现出稳定性能,代码将在论文发表后发布。
英文摘要
Infrared-visible object detection (IVOD) integrates complementary evidence from visible and infrared sensors for reliable perception in challenging scenes. In practice, sensors may fail or drop frames, leaving one modality unavailable or intermittent. Existing methods for IVOD assume both modalities are always present, and fixed fusion collapses when one stream is missing. Furthermore, it remains a critical challenge to reliably estimate semantic correlation across heterogeneous modalities, especially under spectral distribution discrepancy. We present FlexibleFusion, a unified and adaptive method that flexibly allocates integration pathways and fusion strength, operating seamlessly across complete and missing-modality regimes. At its core, the Modality-Aware Experts Collaboration (MAEC) mechanism selectively activates and aggregates cross-modal or intra-modal expert pathways. It allows cross-modal fusion when full modalities are available and falls back to self-fusion under missing conditions. Additionally, we design Residual Self-Paced Entropic Optimal Transport (RSPEOT) to align heterogeneous feature distributions from a transport perspective. Instead of relying on the fixed sparsity coefficient in standard entropic optimal transport (EOT), RSPEOT introduces a residual-driven self-paced update that prioritizes reliable matches and progressively refines harder ones. This design alleviates the additional optimization burden of standard EOT while preserving reliable semantic alignment. Comprehensive experiments under complete and missing-modality protocols show consistent performance across arbitrary modality configurations. Code will be released upon publication.