arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TRIM-ReID:面向多模态目标重识别的去重令牌缩减与模态对齐交互

TRIM-ReID: Duplication-Aware Token Reduction and Modality-Aligned Interaction for Multi-Modal Object Re-Identification

Wanke Xia, Ruiding Zhu, Xingguo Xu, Zhengbo Zhang, Dongxia Liu, Yuan Jin, Taojie Zhu, Yiting Zhao, Yihang Ding

arXiv 2610.04361首次发表:更新:

发表机构

Tsinghua University; Anhui University; Dalian University of Technology; CASIA; SJTU(清华大学; 安徽大学; 大连理工大学; 中国科学院自动化研究所; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TRIM-ReID通过密集身份表示、令牌多样性挖掘和模态关系交互,解决多模态重识别中的令牌冗余与跨模态对齐问题,在三个基准上达到最优性能。

AI 中文摘要

多模态目标重识别利用互补的RGB、近红外(NIR)和热红外(TIR)观测来检索目标对象。然而,现有方法通常采用针对全局图像-文本对齐优化的视觉编码器,并使用学习到的重要性分数来选择令牌。此类设计未能保留细粒度的身份线索,也未明确考虑令牌冗余,导致局部证据代表性不足以及重复令牌引发噪声大且成本高的跨模态交互。为解决这一不足,我们提出TRIM-ReID,一个紧凑框架,统一了密集特征提取、模态内令牌缩减和模态间对齐交互。具体而言,语义丰富且空间连贯的补丁特征由密集身份表示(DIR)提取,其利用DINOv3保留细粒度身份信息。随后,我们引入令牌多样性挖掘(TDM)来识别互补的局部证据,并通过抑制重复补丁同时保留信息多样性来构建紧凑的模态特定令牌集。保留的令牌随后通过模态关系交互(MRI)融合,以实现跨模态的有效信息交换,同时三角对齐损失显式正则化其联合关系,以在独立令牌选择下维持跨模态语义一致性。在RGBNT201、RGBNT100和MSVR310上的大量实验表明,TRIM-ReID实现了最先进的性能。

英文摘要

Multi-modal object re-identification exploits complementary RGB, near-infrared (NIR), and thermal-infrared (TIR) observations to retrieve target objects. However, existing methods commonly employ visual encoders optimized for global image-text alignment and select tokens using learned importance scores. Such designs fail to preserve fine-grained identity cues or explicitly account for token redundancy, resulting in underrepresented local evidence and duplicated tokens that lead to noisy and costly cross-modal interaction. To address this gap, we propose TRIM-ReID, a compact framework that unifies dense feature extraction, intra-modal token reduction, and inter-modal aligned interaction. Specifically, semantically rich and spatially coherent patch features are extracted by Dense Identity Representation (DIR), which leverages DINOv3 to preserve fine-grained identity information. We then introduce Token Diversity Mining (TDM) to identify complementary local evidence and construct compact modality-specific token sets by suppressing repetitive patches while preserving informative diversity. Retained tokens are subsequently fused by Modal Relational Interaction (MRI) to enable effective information exchange across modalities, while a triangular alignment loss explicitly regularizes their joint relationships to maintain cross-modal semantic consistency under independent token selection. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 demonstrate that TRIM-ReID achieves state-of-the-art performance.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑