缓解严重跨模态错位:通过显式特征域仿射配准实现端到端可见光-红外目标检测
Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration
浏览论文内容
中文总结 AI 辅助
针对可见光-红外目标检测的严重跨模态错位问题,提出JFRDet网络,含CMAA、IGCF、AQCG模块,构建DVMA基准,在该基准上达到69.7% mAP₅₀的SOTA性能。
中文摘要 AI 辅助
可见光-红外目标检测依赖互补的RGB(可见光)和热成像线索,但其性能常因跨模态空间错位而下降。现有多数方法依靠隐式特征适配处理轻微错位场景,而大偏移几何差异仍未得到充分解决。本文提出联合特征域配准与检测网络(JFRDet),这是一种专为严重跨模态几何错位设计的端到端可见光-红外定向目标检测器。JFRDet引入跨模态仿射对齐(CMAA)模块,以估计图像级仿射变换实现显式多级特征对齐。考虑到光照变化直接影响RGB线索的可靠性,光照引导互补融合(IGCF)模块可在不同光照条件下自适应利用模态可靠性进行跨模态融合。此外,对齐质量一致性门控(AQCG)策略通过根据对齐可靠性和梯度一致性调节检测监督,稳定联合优化。我们进一步构建了无人机车辆错位数据集(DVMA),用于评估严重跨模态几何错位下的可见光-红外定向目标检测性能。所提出的JFRDet在DVMA上达到69.7%的mAP₅₀,代表了当前最优(SOTA)性能,代码和数据集将在GitHub上发布。
英文摘要
Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently addressed. In this paper, we propose a Joint Feature-domain Registration and Detection network (JFRDet), an end-to-end visible-infrared oriented object detector tailored for severely cross-modal geometric discrepancies. JFRDet introduces a Cross-Modal Affine Alignment (CMAA) module to estimate an image-level affine transformation for explicit multi-level feature alignment. Note that illumination changes directly affect the reliability of RGB cues, an Illumination-Guided Complementary Fusion (IGCF) module adaptively exploits modality reliability under varying illumination conditions for cross-modal fusion. Then, an Alignment Quality-Consistency Gating (AQCG) strategy stabilizes joint optimization by modulating detection supervision according to alignment reliability and gradient consistency. We further construct DroneVehicle Misaligned (DVMA), a benchmark for evaluating visible-infrared oriented object detection under severe cross-modal geometric misalignment. The proposed JFRDet achieves 69.7\% $\mathrm{mAP}_{50}$ on DVMA, which represents state-of-the-art (SOTA) performance. The code and dataset will be available on GitHub.
发表机构
- Beijing University of Technology(北京工业大学)
- Central South University(中南大学)
- Beijing Electronics Science & Technology Institute(北京电子科学技术研究所)
- China Aerospace Science & Industry Corporation(中国航天科技集团)
- Beijing Institute of Technology(北京理工大学)
- Information Support Force Engineering University(信息支援部队工程大学)
- Trunk Technology (Beijing) Co., Ltd.(trunk技术(北京)有限公司)
机构由 AI 辅助整理,请以论文原文为准。