arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于无监督可见光-红外行人重识别的结构-语义交互学习

Structural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-Identification

Moyao Tian, Shijia Liu, Yan Yang, Xin Yuan, Minshi Chen, Wei Wang, Xiao Wang

arXiv 2607.15220首次发表:更新:

发表机构

School of Computer Science and Technology, Wuhan University of Science and Technology; Hubei Province Key Laboratory of Intelligent Information Processing and Real-Time Industrial System, Wuhan University of Science and Technology; Shenyang Institute of Automation, Chinese Academy of Sciences; China University of Chinese Academy of Sciences(武汉科技大学计算机科学与技术学院; 武汉科技大学智能信息处理与实时工业系统湖北省重点实验室; 中国科学院沈阳自动化研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无监督可见光-红外行人重识别的挑战,提出结构-语义交互学习框架SSRL。通过细粒度结构解耦和闭环语义校准机制,实现结构与语义学习的交互,有效过滤伪标签噪声,在相关数据集上表现优于现有方法。

AI 中文摘要

无监督可见光-红外行人重识别(USVI-ReID)因模态差距大且缺乏跨模态身份标注而具有挑战性。渐进关联范式虽被提出以逐步弥合差距,但存在依赖模糊全局表示和伪标签噪声开环传播不受控这两个关键瓶颈。为解决这些问题,我们提出结构-语义交互学习(SSRL)框架,将开环关联转变为自校正闭环系统。结构上引入细粒度结构解耦提取有判别力的身体部位基元作为可靠空间锚点,语义上设计闭环语义校准机制在每个epoch重建共享语义原型并反馈到训练循环,有效过滤伪标签噪声。通过结构和语义学习的交互,SSRL实现了强大的跨模态表示。大量实验表明SSRL在SYSU-MM01和RegDB上优于现有USVI-ReID方法,在RegDB上甚至超过了一些有监督的对应方法。

英文摘要

Unsupervised visible-infrared person re-identification (USVI-ReID) is challenging due to the large modality gap and the lack of cross-modal identity annotations. Progressive association paradigms have been proposed to gradually bridge the gap, but they suffer from two critical bottlenecks: reliance on ambiguous global representations and unchecked propagation of pseudo-label noise in an open-loop manner. To address these issues, we propose Structural-Semantic Reciprocal Learning (SSRL), a framework that transforms open-loop association into a self-correcting closed-loop system. Structurally, we introduce Fine-grained Structural Decoupling (FSD) to extract discriminative body-part primitives as reliable spatial anchors, complementing ambiguous holistic silhouettes with spatially consistent structural details. Semantically, we design a Closed-loop Semantic Calibration (CSC) mechanism that reconstructs shared semantic prototypes at each epoch and feeds them back into the training loop, effectively filtering pseudo-label noise before the next clustering cycle. Through the reciprocal interaction between structural and semantic learning, SSRL achieves robust cross-modal representation. Extensive experiments demonstrate the competitive performance of SSRL against state-of-the-art USVI-ReID methods on both SYSU-MM01 and RegDB, notably surpassing several supervised counterparts on RegDB.

CommentsAccepted by PRCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑