arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式检索用于无监督文本行人搜索

Generative Retrieval for Unsupervised Text-Based Person Search

Mang Ye, Yucheng Ji, Yang Bai, Min Cao, Siyuan Chai, Bo Du, Min Zhang

arXiv 2609.12965首次发表:更新:

发表机构

Soochow University; Wuhan University; Zhipu AI; Harbin Institute of Technology, Shenzhen(苏州大学; 武汉大学; 智谱AI; 哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出GTR+两阶段生成-检索框架,通过分层描述生成与自适应置信度加权学习实现无监督文本行人搜索,并贡献LargeFine-Person数据集,在多个基准上验证了有效性。

AI 中文摘要

文本行人搜索(TBPS)旨在根据给定的自然语言描述,从大型图像库中检索目标人物的图像。现有方法大多依赖于人工标注的图像-文本对进行监督学习。本文探索了无监督TBPS,仅使用未标注图像。我们提出GTR+,一个两阶段的生成-检索框架。在生成阶段,我们引入了一个分层描述生成框架,通过三层顺序过程生成细粒度且风格多样的文本描述。基础层利用自动问答机制生成基本视觉属性描述;中间层通过样本间对比机制增强细粒度细节;高级层通过风格化扩展机制进一步丰富文本多样性。在检索阶段,为减轻噪声伪文本的影响,我们开发了一个自适应置信度加权检索学习框架。我们使用高斯混合模型将图像-文本对建模为干净或噪声,并通过实时图像-文本相似度和来自前一阶段的静态文本生成概率进行校准,在训练过程中产生自适应样本权重。此外,我们还贡献了LargeFine-Person,一个具有高质量、细粒度和多样化文本标注的大规模TBPS数据集,为无监督设置下的实际且可泛化的TBPS预训练基准提供了支持。在多个TBPS基准上的实验证明了GTR+和LargeFine-Person的有效性和泛化性。代码可在以下网址获取:this https URL。

英文摘要

Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on a given natural language description. Most existing methods rely on supervised learning with manually annotated image-text pairs. In this paper, we explore unsupervised TBPS, with only unlabeled images. We propose GTR+, a two-stage generation-then-retrieval framework. In the generation stage, we introduce a tiered description generation framework designed to produce fine-grained and stylistically diverse textual descriptions through a three-tier sequential process. The base tier leverages an automated question-and-answer mechanism to generate basic visual attribute descriptions; the intermediate tier enhances fine-grained detail using an inter-sample contrastive mechanism; the advanced tier further enriches textual diversity via a stylized expansion mechanism. In the retrieval stage, to mitigate the impact of noisy pseudo texts, we develop an adaptive confidence-weighted retrieval learning framework. We model image-text pairs as clean or noisy using a Gaussian Mixture Model, calibrated by real-time image-text similarity and static text generation probability from the prior stage, yielding adaptive sample weights during training. Beyond that, we also contribute LargeFine-Person, a large-scale TBPS dataset with high-quality, fine-grained, and diverse textual annotations, enabling a practical and generalizable TBPS pre-training benchmark under unsupervised setting. Experiments on multiple TBPS benchmarks demonstrate the effectiveness and generalization of both GTR+ and LargeFine-Person. Code is available at: https://github.com/Flame-Chasers/GTR.

Comments17 pages, 10 figures. Accepted for publication in IEEE Transactions on Pattern Analysis and Machine Intelligence

Journal refIEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1-17, 2026

DOI:10.1109/TPAMI.2026.3713379

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑