arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10497cs.CV

SapiensID 2.0:将人类识别基础模型与人类感知对齐

SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception

Yiyang Su, Jie Zhu, Feng Liu, Anil K. Jain, Xiaoming Liu

AI总结:

该研究针对现有人类识别模型与人类感知脱节的问题,提出兼具语义与时间感知的SapiensID 2.0框架,通过多项技术实现最优识别性能。

AI中文摘要:

尽管基础模型已在多模态人类识别领域取得显著进展,但它们主要依赖静态几何特征提取,这种方法与人类感知存在根本差异。因此,当前模型常存在“语义盲区”,过度拟合瞬时噪声而无法利用不变的软生物特征,且难以捕捉时间运动特征。为弥合这一差距,我们提出SapiensID 2.0,这是一个兼具语义与时间感知的人类识别框架。为解决软生物特征标注缺失的问题,我们将多模态大语言模型(MLLMs)的零样本语义知识迁移至判别式嵌入空间;通过不变特征对齐(ITA)解决这些空间间的维度不匹配问题,以提炼核心持久特征;通过瞬时噪声解耦(TND)解耦衣物等干扰因素。此外,我们设计了运动语义注意力头(K-SAH),将空间注意力扩展至时间窗口,通过跟踪语义块随时间的变化,在无需大规模视频数据集的情况下捕捉丰富的运动特征。大量实验表明,SapiensID 2.0在基于图像和视频的行人重识别及步态识别任务中均达到了当前最优性能,同时保持了稳健的人脸识别能力。

英文摘要:

While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fundamentally diverges from human perception. Consequently, current models often suffer from "semantic blindness," overfitting to transient noise while failing to leverage invariant soft biometrics, and struggle to capture temporal motion signatures. To bridge this gap, we propose SapiensID 2.0, a human recognition framework enriched with both semantic and temporal awareness. To overcome the lack of soft-biometric annotations, we transfer zero-shot semantic knowledge from Multimodal Large Language Models (MLLMs) into a discriminative embedding space. We resolve the dimensional mismatch between these spaces using Invariant Trait Alignment (ITA) to distill core persistent traits, and Transient Noise Disentanglement (TND) to decouple artifacts like clothing. Furthermore, we design a Kinematic Semantic Attention Head (K-SAH) that extends spatial attention across temporal windows. By tracking semantic patches over time, K-SAH captures rich kinematic signatures without requiring large-scale video datasets. Extensive experiments demonstrate that SapiensID 2.0 achieves state-of-the-art performance across image- and video-based person re-identification and gait recognition, while maintaining robust face recognition capabilities.

↑