发表机构
Dalian University of Technology; University of California, San Francisco; Nanjing University of Science and Technology(大连理工大学; 加利福尼亚大学旧金山分校; 南京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对多模态物体重识别的语义先验利用不足与全局-局部协同缺失问题,提出含TSI、MGLM、HMF的双语义引导与全局-局部互调制框架,经多基准实验验证有效。
AI 中文摘要
多模态物体重识别(ReID)旨在利用跨模态互补信息检索目标实例,但现有方法存在两大挑战:一是难以利用对齐良好且可靠的语义先验,易受背景杂波和跨模态错位影响;二是通常依赖整体特征建模,忽略全局与局部表征的协同作用。为克服这些局限,本文提出一种融合双语义引导与全局-局部互调制的鲁棒多模态ReID框架,主要包含三个关键组件:文本语义注入器(TSI)、掩码全局-局部调制器(MGLM)和分层门控混合专家融合器(HMF)。TSI通过将干净连贯的文本特征整合到视觉令牌中增强语义感知;MGLM借助软掩码与全局上下文的联合引导实现部件感知的跨模态交互,提升细粒度特征对齐;HMF在局部语义监督下自适应聚合多光谱特征,生成具有判别力且鲁棒的表征。在三个多模态ReID基准上开展的大量实验验证了该方法的有效性,代码将在论文接收后公开于指定URL。
英文摘要
Multi-modal object Re-Identification (ReID) aims to retrieve target instances by leveraging complementary information across modalities. However, existing methods suffer from two challenges. First, they often fail to exploit well-aligned and reliable semantic priors, making them vulnerable to background clutter and cross-modal misalignment. On the other hand, they typically rely on holistic feature modeling, overlooking the synergy between global and local representations. To overcome these limitations, we propose a robust multi-modal ReID framework with dual semantic guidance and global-local mutual modulation, which mainly consists of three key components, namely the Text-Semantic Injector (TSI), the Masked Global-Local Modulator (MGLM), and the Hierarchical MoE Fusion (HMF). The TSI enhances semantic awareness by integrating clean and coherent textual features into visual tokens. The MGLM enables part-aware cross-modal interaction through joint guidance from soft masks and global context, improving fine-grained feature alignment. Finally, the HMF adaptively aggregates multi-spectral features under local semantic supervision, yielding discriminative and robust representations. Extensive experiments on three multi-modal ReID benchmarks demonstrate the effectiveness of the proposed method. The code will be made publicly available at https://github.com/zw-absin/DSGM upon acceptance.
CommentsAccepted by IEEE TCSVT 2026. The version of record may differ slightly