AI 中文总结
研究开放世界场景下可见-红外模态不完整行人重识别问题,提出模态自适应匹配Transformer(MAMT),通过散度Transformer模块和共享Transformer模块提取特征,经模态自适应匹配模块动态融合,在新构建基准上实验,证明方法有效且具适应性。
AI 中文摘要
可见-红外行人重识别(VI-ReID)在封闭世界假设下运行,查询和图库来自异构模态。但在开放世界场景中,两者可能包含同质和异质模态图像。基于异构模态检索范式的VI-ReID方法面临三个可信度挑战:高同质模态相似性导致的匹配冲突、模态不确定性的干扰以及未知模态组合引起的鲁棒性下降。为应对这些挑战,我们将可见-红外模态不完整重识别(VIMI-ReID)任务形式化,重新组织现有数据集构建SYSU-VIMI和RegDB-VIMI基准。现有VI-ReID方法在VIMI-ReID中性能显著下降,我们提出模态自适应匹配Transformer(MAMT),它采用散度Transformer模块(DTM)和共享Transformer模块(STM)分别提取模态特定和模态共享特征。DTM由散度损失引导,用模态风格信息丰富模态特定特征以增强同模态内的可辨别性。模态自适应匹配模块(MAM)根据查询-图库模态关系动态融合特征,在任意和不确定模态条件下实现稳定匹配。在VIMI基准上的大量实验证明了MAMT的有效性和适应性。
英文摘要
Visible-Infrared Person Re-Identification (VI-ReID) operates under a closed-world assumption, where queries and galleries are from heterogeneous modalities. However, in open-world scenarios, both sets are likely to contain homogeneous and heterogeneous modality images. A query may consist of visible-only, infrared-only, or mixed-modality images, while galleries present multi-modal images over long-term collection. Under these conditions, VI-ReID methods, built on a heterogeneous-modality retrieval paradigm, suffer from three trustworthiness challenges: matching conflicts due to high homogeneous-modality similarity, interference from modality uncertainty, and robustness degradation induced by unknown modality combinations. They fail to meet the requirements of trustworthy visual recognition in reliability, consistency, and dynamic adaptability. To address these challenges, we formalize the Visible-Infrared Modality-Incomplete Re-Identification (VIMI-ReID) task. We reorganize existing datasets to construct the SYSU-VIMI and RegDB-VIMI benchmarks. The unpredictable modality combinations and inherent similarity of homogeneous-modality samples in VIMI-ReID cause a significant performance drop in existing VI-ReID methods. We propose the Modality Adaptive Matching Transformer (MAMT). It employs a Divergence Transformer Module (DTM) and a Shared Transformer Module (STM) to extract modality-specific and modality-shared features, respectively. Guided by a divergence loss, the DTM enriches modality-specific features with modality-style information to enhance discriminability within the same modality. A Modality Adaptive Matching Module (MAM) dynamically fuses features according to the query-gallery modality relationship, enabling stable matching under arbitrary and uncertain modality conditions. Extensive experiments on the VIMI benchmarks demonstrate the effectiveness and adaptability of MAMT.
Comments18 pages, 7 figures