arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiffReID:用于目标重识别的判别式扩散模型

DiffReID: Discriminative Diffusion Model for Object Re-Identification

Yingquan Wang, Pingping Zhang, Dong Wang, Huchuan Lu

arXiv 2609.36894首次发表:更新:

发表机构

Dalian University of Technology(大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DiffReID框架,利用判别式扩散模型学习身份感知分布并生成身份不变特征,通过VNG、LWD和MEC模块提升目标重识别的泛化与判别能力,在五个基准上优于多数现有方法。

AI 中文摘要

作为一项基础的图像处理任务,目标重识别(ReID)旨在跨不重叠的摄像头检索目标。近年来,随着深度学习的发展,目标重识别取得了显著进展。然而,由于ReID数据集规模有限且多样性不足,大多数现有方法在泛化方面表现不佳。同时,当前模型倾向于提取语义模式,而非学习身份感知的特征分布。为解决这些问题,我们提出了一种名为\ extbf{DiffReID}的新型特征学习框架,用于目标重识别。该框架利用判别式扩散模型逐步学习身份感知分布并生成身份不变特征。具体而言,借助对比语言-图像预训练(CLIP)模型,我们首先通过提示调优获得身份感知的文本特征。然后,我们提出了一种视觉引导噪声生成器(VNG)来初始化概率噪声并逐步破坏身份感知的文本特征。之后,我们将视觉特征作为条件,并提出了一种轻量级去噪器(LWD)来逐步对损坏的文本特征进行去噪,以实现身份感知分布学习。为了获得判别性特征,我们进一步从随机采样的视觉引导噪声中生成身份不变的引导特征。最后,我们提出了一种互增强约束(MEC),以促进视觉特征与引导特征之间的相互学习,从而增强表示的鲁棒性和判别性。在五个目标重识别基准上的大量实验表明,我们的方法优于大多数最先进的方法。源代码可在https://这个URL获取。

英文摘要

As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing methods suffer from generalization due to the limited size and diversity of ReID datasets. Meanwhile, current models tend to focus on extracting semantic patterns rather than learning identity-aware feature distributions. To address these issues, we propose a novel feature learning framework named \textbf{DiffReID} for object ReID. It leverages a discriminative diffusion model to gradually learn identity-aware distributions and generate identity-invariant features. More specifically, with the Contrastive Language-Image Pre-training (CLIP) model, we first obtain identity-aware text features by prompt tuning. Then, we propose a Vision-guided Noise Generator (VNG) to initialize probabilistic noises and gradually corrupt identity-aware text features. Afterwards, we take visual features as conditions and propose a Light Weight Denoiser (LWD) to denoise the corrupted text features step-by-step for identity-aware distribution learning. To obtain discriminative features, we further generate identity-invariant guided features from randomly sampling visual-guided noises. Finally, we propose a Mutual Enhancement Constraint (MEC) to facilitate mutual learning between visual features and guided features to enhance the representation robustness and discrimination. Extensive experiments on five object ReID benchmarks demonstrate that our method shows better results than most state-of-the-art methods. The source code is available at https://github.com/AWangYQ/DiffReID.

CommentsAccepted by TIP2026. More modifications can be performed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑