arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06205cs.CV

CFGPNet:用于多光谱目标检测的基于交叉注意力的融合梯度编程网络框架

CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection

Nima Hatami, Karim Faez, Saeed Sharifian, Hamidreza Amindavar

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有RGB-T目标检测方法的跨模态交互不足等问题,提出CFGPNet框架,采用改进GELAN骨干、CrossCEA模块等设计,在五个多光谱基准数据集上取得优异性能,实现精度与效率的良好权衡。

中文摘要 AI 辅助

RGB-T目标检测利用可见光和红外图像的互补优势,支持在低光照、恶劣天气和复杂多尺度环境下的鲁棒感知。然而,现有方法仍存在跨模态交互不足、因模态分布差异导致的融合不稳定,以及基于注意力的重型架构计算成本高的问题。为解决这些问题,本文提出用于多光谱目标检测的CFGPNet,即基于交叉注意力的融合梯度编程网络框架。CFGPNet采用改进的GELAN骨干网络,结合RepViT风格的重参数化模块,在保持计算效率的同时增强特征表示;引入跨计算高效注意力(CrossCEA)模块,强化可见光分支与热红外分支之间的跨模态特征交互,减少冗余信息传递;为生成紧凑且具有判别性的融合表示,采用注意力选择与聚合融合(ASAF)网络,将密集特征聚合与基于选择性注意力的强调相结合;此外,在每个CFGPNet变体中集成可编程梯度辅助分支,以改善梯度传递和优化质量。在FLIR、M3FD、LLVIP、VEDAI和MFAD五个公开多光谱基准数据集上的实验表明,CFGPNet在不同场景、目标尺度和模态平衡条件下均实现了稳定且优异的性能,具体而言,在FLIR数据集上达到80.7%的mAP50和45.0%的mAP50:95,在M3FD数据集上为89.9%/63.4%,在LLVIP数据集上为97.8%/68.9%,在VEDAI数据集上为83.3%/56.9%,在MFAD数据集上为83.4%/61.8%。这些结果表明,CFGPNet是一种有效且实用的解决方案,在三种模型尺度下均能实现精度与效率的良好权衡,其代码、数据和微调模型可在指定URL获取。

英文摘要

Multispectral object detection combines visible and thermal imagery to improve perception under challenging illumination and environmental conditions. However, differences in modality appearance and reliability can introduce redundant or conflicting responses, limiting the use of complementary information. Complex fusion mechanisms further increase computational cost, creating a persistent trade-off between detection accuracy and efficiency. To address these challenges, CFGPNet is proposed, a cross-attention-based fused gradient programmed network. The framework incorporates re-parameterized RepViT blocks into the YOLOv9 architecture to strengthen spatial and channel representations while maintaining efficient feature extraction. Cross Computation Efficient Attention (CrossCEA) exchanges spatial attention maps between modalities at multiple detection scales, allowing each stream to emphasize regions supported by the other while preserving modality-specific information. Attention Selection and Aggregation Fusion (ASAF) combines dense feature aggregation with selection of the strongest responses from multiple attention branches to form compact, discriminative fused representations. A programmable gradient information pathway provides auxiliary supervision during training to improve feature learning. This pathway is removed after training, adding no parameters or operations at inference. Experiments on FLIR, M3FD, LLVIP, VEDAI, and MFAD demonstrate favorable accuracy-efficiency trade-offs across three model scales, with the smallest variant requiring 15.3 million parameters and 56.9 GFLOPs. The code is available at https://github.com/NimaHatami99/CFGPNet.

发表机构

  • Amirkabir University of Technology(阿米尔卡比尔理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑