用于稳健光学-SAR目标检测的边界对齐贡献路由
Boundary-Aligned Contribution Routing for Robust Optical--SAR Object Detection
- Tianjin University(天津大学)
- School of Electrical and Information Engineering, Tianjin University(天津大学电气与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对光学-SAR融合目标检测中的负跨模态迁移问题,提出边界对齐贡献路由方法,在M4-SAR和SpaceNet6-OTD实验中提升了检测性能并降低了负迁移率。
AI中文摘要:
光学图像提供丰富的外观线索,而合成孔径雷达(SAR)的观测结果对光照和天气敏感度较低,这使得光学-SAR融合在遥感目标检测中颇具吸引力。然而,多模态的存在并不一定保证融合有益:空间、时间和语义对应关系的不完善会使原本完整的数据流产生条件性危害,并引发负跨模态迁移。我们从特定任务的任务效用视角处理这一问题,仅利用检测监督学习任务条件下的贡献路由。所提出的融合边界对齐路由在首次学习跨模态特征值混合操作之前,调节每个模态的贡献。对于频繁浅层交互的架构,特征路由器在输入附近执行跨条件、组可寻址调制;对于双骨干架构,双统计语义路由器在后期语义融合前,从特定模态的平均和最大统计量预测流级贡献权重。这些路由器无需显式效用监督、质量标签、重建或蒸馏。在M4-SAR和SpaceNet6-OTD上的实验涵盖标称全输入、受控对应偏移、缺失模态以及四种非零模态损坏场景。在报告的干净训练对照中,路由使全输入的mAP₅₀提升0.5至5.9个百分点;相对于相应的模态丢弃基线,它使缺失模态的mAP₅₀提升7.6至41.6个百分点,并将负迁移率降低多达12.7个百分点。学习到的路由权重与特定模态留一法效用之间的斯皮尔曼相关系数范围为0.45至0.66,支持路由系数的任务效用解释。
英文摘要:
Optical imagery provides rich appearance cues, whereas synthetic aperture radar (SAR) offers observations that are less sensitive to illumination and weather, making optical--SAR fusion attractive for remote-sensing object detection. However, the presence of multiple modalities does not guarantee beneficial fusion: imperfect spatial, temporal, and semantic correspondence can make an otherwise intact stream conditionally harmful and induce negative cross-modal transfer. We handle this issue through a model-specific task-utility perspective and learn task-conditioned contribution routing using detection supervision alone. The proposed fusion-boundary-aligned routing regulates each modality's contribution before the first learned cross-modal feature-value mixing operation. For architectures with frequent shallow interaction, a Feature Router performs cross-conditioned, group-addressable modulation near the input; for dual-backbone architectures, a Dual-Statistic Semantic Router predicts stream-level contribution weights from modality-specific average and maximum statistics before late semantic fusion. The routers require no explicit utility supervision, quality labels, reconstruction, or distillation. Experiments on M4-SAR and SpaceNet6-OTD cover nominal full inputs, controlled correspondence shifts, missing modalities, and four nonzero modality-corruption scenarios. Across the reported clean-training controls, routing improves full-input $\text{mAP}_{50}$ by 0.5--5.9 points. Relative to the corresponding modality-dropout baselines, it raises missing-modality $\text{mAP}_{50}$ by 7.6--41.6 points and reduces the negative-transfer rate by up to 12.7 percentage points. Spearman correlations between the learned routing weights and model-specific leave-one-modality-out utility range from 0.45 to 0.66, supporting the task-utility interpretation of the routing coefficients.