MMGait:跨异构模态步态识别的基准测试与统一
MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities
浏览论文内容
中文总结 AI 辅助
提出MMGait多传感器步态基准,并设计OmniGait++统一框架,在共享身份空间内实现单模态、跨模态和多模态识别,验证了异构模态下统一识别的可行性。
中文摘要 AI 辅助
步态识别通常使用RGB视频或其衍生的轮廓和姿态进行研究。然而,人类行走会产生异构的光度、几何和运动线索,这些无法通过以RGB为中心的基准进行系统检验。我们提出了MMGait,一个大规模多传感器基准,将可见光、红外、深度、LiDAR和雷达观测纳入序列级对应关系。它提供了涵盖外观、轮廓、几何、运动和身体结构的多样化模态。在共享的冒名顶替者增强协议下,我们评估了单模态识别、通过定向检索的跨模态识别以及使用任务特定专家的多模态识别。在各种设置中,模态排名随探测条件而变化,跨模态对齐仍然困难,而融合通常提供互补增益。这一分析揭示了一个可扩展性问题:单个模态、模态对和融合配置通常由单独训练的专家处理。我们提出了全模态步态识别,在共享身份空间内统一单模态、跨模态和多模态识别。OmniGait++使用模态特定的前端,后接共享身份编码器,以保留模态依赖线索,同时学习可比较的身份描述符。锚引导融合模块聚合不同大小的模态子集,无需帧级同步。联合训练的检查点覆盖所有三种识别设置,并适应不同组成和基数的模态子集。实验表明,OmniGait++在许多共享设置中与任务特定专家保持竞争力,并扩展到固定配对模型无法实现的更高基数融合。结果确立了MMGait作为异构步态感知的通用测试平台,并证明了在模态可用性变化下统一识别的可行性。
英文摘要
Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmarks. We present MMGait, a large-scale multi-sensor benchmark that brings visible, infrared, depth, LiDAR, and radar observations into sequence-level correspondence. It provides diverse modalities spanning appearance, contours, geometry, motion, and body structure. Under a shared impostor-augmented protocol, we evaluate single-modal recognition, cross-modal recognition via directed retrieval, and multi-modal recognition using task-specific experts. Across settings, modality rankings vary with probe conditions, cross-modal alignment remains difficult, and fusion often provides complementary gains. This analysis exposes a scalability problem: individual modalities, modality pairs, and fusion configurations are typically handled by separately trained experts. We formulate Omni-Modal Gait Recognition, which unifies single-modal, cross-modal, and multi-modal recognition within a shared identity space. OmniGait++ uses modality-specific front ends followed by a shared identity encoder to preserve modality-dependent cues while learning comparable identity descriptors. An anchor-guided fusion module aggregates modality subsets of varying size without frame-level synchronization. A jointly trained checkpoint covers all three recognition settings and accommodates modality subsets of different compositions and cardinalities. Experiments show OmniGait++ remains competitive with task-specific experts in many shared settings and extends to higher-cardinality fusion unavailable to fixed-pair models. The results establish MMGait as a common testbed for heterogeneous gait sensing and demonstrate the feasibility of unified recognition under varying modality availability.
发表机构
- Beijing Normal University(北京师范大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- Watrix Technology Limited Co. Ltd(眼神科技)
机构由 AI 辅助整理,请以论文原文为准。