arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

两阶段多视角步态识别与重嵌入网络

Two-Stage Multi-View Gait Recognition with a Re-Embedding Network

Long Hoang Le, Trung Thanh Ngo

arXiv 2609.32244首次发表:更新:

AI 中文总结

提出TFTR两阶段框架,通过孪生网络和Transformer重嵌入,实现多视角步态识别,在OU-MVLP和CASIA-B上取得高准确率。

AI 中文摘要

步态识别因严重过拟合和单阶段方法中常见的刚性视角约束而始终具有挑战性。我们提出一个两阶段框架,称为“先平移后推理”(Translate-First-Then-Reason, TFTR),以解决这些问题。在第一阶段,一个带有三元组损失的浅层孪生卷积网络将步态能量图(GEIs)映射到128维的视角特定嵌入空间。在第二阶段,这些逐视角嵌入被视为令牌,并由一个12层Transformer编码器处理,该编码器将它们重新投影到一个具有改进余弦可分性的新空间。这种设计能够在推理时灵活融合任意数量的视角,克服了先前方法的固定输入限制。在OU-MVLP数据集(6,000个受试者)上训练,并在未见过的CASIA-B上评估,涵盖正常、携带包和穿外套条件,我们的流程在OU-MVLP上实现了96.91%的单视角和99.49%的三视角准确率,并在CASIA-B上以三个视角达到了100%的准确率。

英文摘要

Gait recognition always remains challenging due to severe overfitting and the rigid view constraints common in single-stage approaches. We propose a two-stage framework, termed Translate-First-Then-Reason (TFTR), to address these issues. In the first stage, a shallow Siamese convolutional network with triplet loss maps Gait Energy Images (GEIs) into a 128-dimensional view-specific embedding space. In the second stage, these per-view embeddings are treated as tokens and processed by a 12-layer Transformer encoder, which re-projects them into a new space with improved cosine separability. This design enables flexible fusion of an arbitrary number of views at inference, overcoming the fixed-input limitations of prior methods. Trained on the OU-MVLP dataset (6,000 subjects) and evaluated on unseen CASIA-B across normal, bag-carrying, and coat-wearing conditions, our pipeline achieves 96.91\% single-view and 99.49\% three-view accuracy on OU-MVLP, and attains 100\% accuracy on CASIA-B with three views.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑