发表机构
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出词汇引导的步态识别范式Gait-World及模型α-Gait,利用视觉-语言模型词汇知识引导步态网络学习,通过词汇关系映射器和细粒度检测器解决模态异质性,在多个数据集上验证了有效性。
AI 中文摘要
什么是步态?基于外观的步态网络将步态视为来自图像的人体形状和运动信息。基于模型的步态网络将步态视为来自点的人体固有结构。然而,这些考虑对于人类真正理解而言仍然模糊。在这项工作中,我们引入了一种新颖的范式——词汇引导的步态识别,称为Gait-World,它试图通过视觉-语言模型(VLMs)利用人类词汇来探索步态概念。尽管VLMs在各种视觉任务中取得了显著进展,但其关于步态模态的认知能力仍然有限。Gait-World的成功要素在于恰当的词汇提示,该范式精心选择步态周期动作作为词汇基础,桥接步态和词汇特征空间,并进一步促进人类对步态的理解。如何提取步态特征?尽管先前的步态网络已取得重大进展,但仅在有限的步态数据库上从步态模态学习,难以学习到适用于实际应用的通用步态特征。因此,我们提出了第一个Gait-World模型,称为{\alpha}-Gait,它利用来自VLMs的词汇知识引导步态网络学习。然而,由于模态的异质性,直接整合词汇和步态特征极具挑战性,因为它们位于不同的嵌入空间。为解决这些问题,{\alpha}-Gait设计了词汇关系映射器和步态细粒度检测器,以在步态空间中映射并建立词汇关系,从而检测相应的步态特征。在CASIA-B、CCPG、SUSTech1K、Gait3D和GREW上的大量实验揭示了VLMs词汇信息在步态领域的潜在价值和研究方向。
英文摘要
What is a gait? Appearance-based gait networks consider a gait as the human shape and motion information from images. Model-based gait networks treat a gait as the human inherent structure from points. However, the considerations remain vague for humans to comprehend truly. In this work, we introduce a novel paradigm Vocabulary-Guided Gait Recognition, dubbed Gait-World, which attempts to explore gait concepts through human vocabularies with Vision-Language Models (VLMs). Although VLMs have achieved the remarkable progress in various vision tasks, the cognitive capability regarding gait modalities remains limited. The success element in Gait-World is the proper vocabulary prompt where this paradigm carefully selects gait cycle actions as Vocabulary Base, bridging the gait and vocabulary feature spaces and further promoting human understanding for the gait. How to extract gait features? Although previous gait networks have made significant progress, learning solely from gait modalities on limited gait databases makes it difficult to learn universal gait features for practicality. Therefore, we propose the first Gait-World model, dubbed α-Gait, which guides the gait network learning with vocabulary knowledge from VLMs. However, due to the heterogeneity of the modalities, directly integrating vocabulary and gait features is highly challenging as they reside in different embedding spaces. To address the issues, α-Gait designs Vocabulary Relation Mapper and Gait Fine grained Detector to map and establish vocabulary relations in the gait space for detecting corresponding gait features. Extensive experiments on CASIA-B, CCPG, SUSTech1K, Gait3D and GREW reveal the potential value and research directions of vocabulary information from VLMs in the gait field.
CommentsAccepted at NeurIPS 2025