发表机构
Anhui Provincial Key Laboratory of Multimodal Cognitive Computation; School of Computer Science and Technology; Anhui University; School of Artificial Intelligence(安徽省多模态认知计算重点实验室; 计算机科学与技术学院; 安徽大学; 人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RGBT视频目标检测中图像对空间未对齐问题,提出双相关超图网络DHNet,通过设计PSAM和DHFM模块捕获互补信息,还构建DVT-VOD1000数据集,实验表明该网络检测精度达最优,数据集和代码将公开。
AI 中文摘要
RGB-Thermal(RGBT)视频目标检测(VOD)因能克服传统基于RGB的VOD在挑战性条件下的局限性而备受关注。然而,RGBT图像对之间通常存在空间未对齐问题。为此,我们提出双相关超图网络(DHNet),通过显式建模连续帧间的时间相关性和跨模态特征的空间相关性来捕获高维互补信息。具体先设计基于补丁的空间对齐模块(PSAM)在局部区域对齐多模态特征,接着引入双超图融合模块(DHFM)构建时间和多模态超图通过双相关学习增强目标可辨别性。此外,目前缺乏大规模、场景多样的基准数据集。我们构建了包含103464对RGBT图像的1000个视频序列的DVT-VOD1000数据集。在VT-VOD50和DVT-VOD1000上的综合实验表明DHNet实现了最优检测精度。数据集和代码将公开支持学术研究。
英文摘要
RGB-Thermal (RGBT) Video Object Detection (VOD) has gained significant attention because of the limitations of conventional RGB-based VOD methods under challenging conditions, such as low light, heavy fog, and adverse weather, etc. However, spatial misalignment commonly exists between RGBT image pairs. To address this, we propose a Dual-Correlation Hypergraph Network (DCHNet) that captures high-dimensional complementary information by explicitly modeling two types of correlations: temporal correlation across consecutive frames and spatial correlation from cross-modal features. Specifically, we first design a Patch-based Spatial Alignment Module (PSAM) to sequentially align the multimodal features at the local region level. Subsequently, we propose a Dual Hypergraph Fusion Module (DHFM), which constructs temporal and multimodal hypergraphs, respectively, to enhance object characteristic through dual-correlation learning. Furthermore, the field currently lacks a large-scale, scene-diverse benchmark dataset for comprehensive evaluation. Therefore, we construct DVT-VOD1000, a large-scale RGBT VOD dataset containing 1,000 video sequences with 103,464 RGBT image pairs. The dataset covers diverse scenarios, including campuses, parks, traffics, rural areas, night scenes, rainy weather, and snowy weather. Comprehensive experiments on VT-VOD50 and our DVT-VOD1000 demonstrate that DCHNet achieves state-of-the-art detection accuracy. The dataset and source code will be made publicly available on https://github.com/tzz-ahu/ to support academic research.