arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCOUT:通过冻结视频特征上的嵌入空间预测实现文本到人物的仿真到现实检索

SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features

Abdarahmane Traoré, Andy Couturier, Éric Hervet

arXiv 2609.19483首次发表:更新:

发表机构

Université de Moncton(蒙克顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SCOUT,利用冻结编码器和嵌入空间预测实现仿真到现实的文本人物检索,通过精度杠杆和校准研究,在AI City Challenge 2026上达到84.25 mAP@10。

AI 中文摘要

在仿真到现实差距(合成训练数据、真实图像库)下的文本到人物检索通常通过代价高昂的微调交叉编码器来解决。我们探究冻结编码器系统是否能与之竞争。我们提出SCOUT,将跨模态检索视为嵌入空间中的预测。一个可训练预测器在双向InfoNCE目标下,将冻结视频编码器的补丁令牌映射到冻结文本编码器的嵌入空间,基础模型中不微调任何编码器。视频编码器为V-JEPA,文本编码器为EmbeddingGemma,预测器从Qwen3.5-0.8B解码器初始化。我们有三项发现。首先,最佳冻结文本编码器只是其几何结构与视频特征最匹配的那个。一个免训练的对齐分数对三个候选文本编码器的排序与它们在我们留出分割上的检索准确率顺序一致(Spearman ρ = 1.0);第四个基于LLM的编码器显示该规则依赖于度量,对邻域重叠分数成立(ρ = 0.8),但对线性探针不成立(ρ = -0.2)。其次,两个针对精度的杠杆——视频编码器的参数高效ExPLoRA适配和基于视觉语言模型的免训练属性分解重排序器——提高了冻结系统原本受限的顶级排名精度,在排行榜上增加了2.2个R@1点。第三,一项局部与公共校准研究解释了哪些干预措施能迁移到真实领域。在AI City Challenge 2026 Track 4上,完整的检索-融合-重排序系统在最终排行榜上达到84.25 mAP@10,而单独提交的单个冻结模型达到60.63。我们训练的组件成本约95个GPU小时。CMP,即数据集作者的微调交叉编码器,训练十六个GPU天,是完整系统的一个融合成员,而非替代方案。代码和注释:此https URL

英文摘要

Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as prediction in embedding space. A trainable predictor maps the patch tokens of a frozen video encoder into the embedding space of a frozen text encoder under a bidirectional InfoNCE objective, and no encoder is fine-tuned in the base model. The video encoder is V-JEPA, the text encoder is EmbeddingGemma, and the predictor is initialized from a Qwen3.5-0.8B decoder. We make three findings. First, the best frozen text encoder is simply the one whose geometry best matches the video features. A training-free alignment score ranks three candidate text encoders in the same order as their retrieval accuracy on our held-out split (Spearman $ρ= 1.0$); a fourth, LLM-based encoder shows the rule is metric-dependent, holding for a neighborhood-overlap score ($ρ= 0.8$) but not for a linear probe ($ρ= -0.2$). Second, two precision-targeted levers, parameter-efficient ExPLoRA adaptation of the video encoder and a training-free attribute-decomposed reranker built on a vision-language model, improve the top-rank precision that otherwise limits the frozen system, adding 2.2 points of leaderboard R@1. Third, a local-versus-public calibration study explains which interventions transfer to the real domain. On AI City Challenge 2026 Track 4 the full retrieve-fuse-rerank system reaches 84.25 mAP@10 on the final leaderboard, while a single frozen model submitted alone reaches 60.63. Our trained components cost about 95 GPU-hours. CMP, the dataset authors' fine-tuned cross-encoder that trains for sixteen GPU-days, is one fusion member of the full system, not an alternative. Code and annotations: https://github.com/abtraore/SCOUT-ECCV

Comments16 pages, 4 figures, 3 tables. Accepted at the ECCV 2026 Workshop on AI City Challenge (Track 4). Code and annotations: https://github.com/abtraore/SCOUT-ECCV

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑