用于通用跨模态行人重识别的双空间模态一致性学习
Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification
- School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
- School of Humanities and Social Sciences, Beihang University(北京航空航天大学人文与社会科学高等研究院)
- School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
- School of Informatics, Xiamen University(厦门大学信息学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DSMCL即插即用框架,联合建模空间与频域一致性,在5个数据集的17项协议上提升跨模态ReID基线性能,适配多样化异构模态场景。
AI中文摘要:
跨模态行人重识别(ReID)旨在检索不同异构成像模态下的同一身份,已在可见光-红外行人ReID、跨模态船舶ReID等场景中得到广泛研究。现有方法通过在空间嵌入空间中学习模态一致性已取得良好性能,但常忽略频域模态差异,尤其是兼具高判别性与模态敏感性的高频表示。此外,多数方法针对特定模态设置设计,限制了其在多样化跨模态场景中的适用性。为解决这些挑战,本文提出双空间模态一致性学习(DSMCL)框架用于通用跨模态ReID。具体而言,DSMCL联合建模空间特征分布一致性与频域判别一致性:空间模态一致性学习(SMCL)分支执行基于高斯的特征对齐,频域感知判别一致性学习(FDCL)策略通过身份感知跨模态对比学习正则化高频表示。通过联合捕捉模态特异性特征与模态共享身份线索,DSMCL学习鲁棒表示并建立可适配多样化异构模态设置的统一框架,且为即插即用框架可便捷集成至现有跨模态ReID架构。在SYSU-MM01、RegDB、LLCM、HOSS-ReID、CMShipReID数据集的17项评估协议上开展的大量实验表明,DSMCL可持续提升多个代表性基线的性能。
英文摘要:
Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-infrared person ReID and cross-modal ship ReID. Existing methods have achieved promising performance by learning modality consistency in the spatial embedding space, yet often overlook frequency-domain modality discrepancy, particularly in high-frequency representations that are both highly discriminative and modality-sensitive. In addition, most approaches are tailored to specific modality settings, limiting their applicability across diverse cross-modal scenarios. To address these challenges, we propose a Dual-Space Modality Consistency Learning (DSMCL) framework for universal cross-modal ReID. Specifically, DSMCL jointly models spatial feature distribution consistency and frequency-domain discriminative consistency. A Spatial Modality Consistency Learning (SMCL) branch performs Gaussian-based feature alignment, while a Frequency-aware Discriminative Consistency Learning (FDCL) strategy regularizes high-frequency representations through identity-aware cross-modal contrastive learning. By jointly capturing modality-specific characteristics and modality-shared identity cues, DSMCL learns robust representations and establishes a unified framework capable of accommodating diverse heterogeneous modality settings. Moreover, DSMCL is a plug-and-play framework that can be readily integrated into existing cross-modal ReID architectures. Extensive experiments on SYSU-MM01, RegDB, LLCM, HOSS-ReID, and CMShipReID across seventeen evaluation protocols show that DSMCL consistently improves multiple representative baselines.