arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26632cs.CV

谁留存,什么改变:基于身份锚定的组合步态检索

Who Remains, What Changes: Identity Anchored Composed Gait Retrieval

Jingchen Fei, Zengbin Wang, Yukun Liu, Muyi Sun, Shibiao Xu, Man Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出组合步态检索(CoGR)新任务,构建了首批步态-语言数据集,提出基于身份锚定的ComposeGait框架,在两个基准上取得最优R@1,为CoGR提供强基线。

中文摘要 AI 辅助

步态识别已取得显著进展,但现有方法仍局限于刚性视觉匹配,往往忽视自然语言指令在交互式检索中的潜力。本文提出组合步态检索(Composed Gait Retrieval, CoGR)这一新任务,其核心是基于参考序列与自然语言修改查询来检索目标步态序列。针对该任务缺乏现有数据集的问题,我们设计了一个由大视觉语言模型(VLMs)驱动的自动标注流水线,构建了首批步态-语言数据集:语言增强版CCPG与语言增强版CASIA-B。在此基础上,我们提出了ComposeGait,这是一个基于身份锚定的组合框架,旨在防止通用组合检索在遵循指令时出现身份漂移(即返回错误人物)的问题。其部分感知身份适配器(Part-aware Identity Adapter, PIA)将多帧、部分感知的身份证据聚合为样本特定的ID令牌。我们将这些ID令牌注入共享Q-Former的两个分支以保留身份,同时在最终检索嵌入中排除ID令牌的输出。联合身份与任务适配的组合检索目标对该空间进行端到端优化。我们在两个基准上对ComposeGait进行评估,结果显示其在对比方法中达到最优的R@1指标,在语言增强版CCPG上为72.38%,在语言增强版CASIA-B上为83.61%。这些结果确立了ComposeGait作为CoGR任务的强基线,相关数据集与代码将公开提供。

英文摘要

Gait recognition has achieved remarkable progress, yet existing methods remain confined to rigid visual matching and often overlook the potential of natural language instructions for interactive retrieval. In this paper, we introduce Composed Gait Retrieval (CoGR), a novel task that retrieves a target gait sequence based on a reference sequence and a natural language modification query. To address the absence of existing datasets for this task, we design an automated annotation pipeline powered by large vision-language models (VLMs) to construct the first gait-language datasets: Language-Augmented CCPG and Language-Augmented CASIA-B. Building on this, we propose ComposeGait, an identity-anchored composition framework designed to prevent the identity drift that arises when generic composed retrieval follows the instruction but returns the wrong person. Its Part-aware Identity Adapter (PIA) aggregates multi-frame, part-aware identity evidence into a sample-specific ID token. We inject the ID tokens into both branches of a shared Q-Former to preserve identity, while excluding the ID-token outputs from the final retrieval embeddings. Joint identity and task-adapted composed-retrieval objectives optimize this space end to end. We evaluate ComposeGait on both benchmarks and show that it achieves the best R@1 among the compared methods, reaching 72.38% on Language-Augmented CCPG and 83.61% on Language-Augmented CASIA-B. These results establish ComposeGait as a strong baseline for CoGR. The datasets and code will be made publicly available.

发表机构

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Huazhong University of Science and Technology(华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑