发表机构
College of Computer Science and Technology, National University of Defense Technology; School of Computer Science and Information Engineering, Hefei University of Technology(国防科技大学计算机科学与技术学院; 合肥工业大学计算机与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MODAL多模态对象重识别框架,通过多模态特征稀疏解耦模块与文本-图像差分滤波模块解决现有方法的特征纠缠与模态缺失性能下降问题,在四个数据集上取得SOTA性能。
AI 中文摘要
多模态对象重识别(Re-ID)旨在通过利用视觉(如RGB、近红外NIR、热红外TIR)和文本模态的互补信息,促进复杂环境下跨摄像头的对象检索。然而,现有方法往往缺乏原则性的特征解耦与连贯的多模态融合,导致表征纠缠,进而引发跨模态冲突、掩盖判别线索,并在模态缺失条件下出现分布偏移。为应对这些挑战,本文提出MODAL,一个基于耦合稀疏编码理论与差分抑制原则的新型多模态对象重识别框架。MODAL的核心组件是多模态特征稀疏解耦模块,该模块以模型驱动的深度展开方式,基于多模态耦合稀疏编码开发,明确将多模态特征分解为单模态特有、双模态共享及三模态共享表征,从而实现更透明、有效的特征解耦。得益于该原则性特征解耦,MODAL通过模态感知子空间激活,仅选择性激活一致共享子空间,自然缓解了不完整模态场景下的性能下降。此外,本文提出文本-图像差分滤波模块,该模块利用粗粒度文本语义自适应抑制解耦视觉表征中与任务无关的响应,进而增强判别信息。在四个数据集上的大量实验表明,MODAL实现了超越现有方法的性能,且具备更优的透明性。
英文摘要
Multi-modal object re-identification (Re-ID) aims to facilitate cross-camera object retrieval in complex environments by leveraging complementary information from visual (e.g., RGB, NIR, TIR) and textual modalities. However, existing approaches often lack principled feature disentanglement and coherent multi-modal integration, leading to entangled representations that introduce cross-modal conflicts, obscure discriminative cues, and suffer distribution shift under modality-missing conditions. To tackle these challenges, we propose MODAL, a novel multi-modal object re-identification framework, grounded in coupled sparse coding theory and differential suppression principles. A core component of MODAL is a Multi-modal Feature Sparse Decoupling module, developed in a model-driven deep unrolling manner based on multi-modal coupled sparse coding. It explicitly decomposes multi-modal features into uni-modal specific, bi-modal and tri-modal shared representations, thereby achieving more transparent and effective feature disentanglement. Benefiting from the principled feature disentanglement, MODAL naturally mitigates performance degradation in incomplete-modality scenarios via a Modality-Aware Subspace Activation that selectively activates only the consistently shared subspaces. Moreover, we propose a Text-Image Differential Filtering module that leverages coarse-grained textual semantics to adaptively suppress task-irrelevant responses in the decoupled visual representations, thereby enhancing discriminative information. Extensive experiments on four datasets demonstrate that MODAL achieves state-of-the-art performance with superior transparency.