arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05242cs.CV

基于层次度量学习的少样本视频识别

Few-Shot Video Recognition via Hierarchical Metric Learning

  • School of Information Science and Engineering, University of Jinan(济南大学信息科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

Jiaxin Zhang, Haoran Gao, Xizhan Gao, Zihao Dong, Tingwei Wang, Sijie Niu

中文总结 AI 辅助

针对少样本动作识别中现有模型泛化能力不足的问题,提出HML-FSAR方法,通过构建完整特征流水线并施加多阶段互补约束,在5个FSAR数据集上验证了方法的有效性。

中文摘要 AI 辅助

少样本动作识别(FSAR)旨在仅用少量标注视频样本识别未见过的动作类别。现有工作通常在网络输出端应用单原型监督,未能充分利用视频中丰富的跨帧全局空间信息;即便现有多级度量方案仅对中间层施加并行原型约束,未沿完整特征流水线进行渐进式监督,导致学习到的类别原型泛化能力有限。受此启发,我们提出一种新颖方法——面向少样本动作识别的层次度量学习(HML-FSAR)。首先,开发空间增强模块以捕获跨帧全局空间表示,结合时间多头注意力(MHA)、异构对齐、时空特征融合及字典学习模块,构建完整的特征处理流水线。其次,将层次度量学习(HML)策略嵌入HML-FSAR,该策略由中心度量、对齐度量、对比度量、字典度量及原型度量组成,从帧级表示到最终类别原型施加渐进式多阶段互补约束,以联合优化特征紧凑性、异构时空对齐、类间判别性及抗噪鲁棒性。所提HML-FSAR方法在5个广泛使用的FSAR数据集上得到验证,实验结果充分证明了其有效性。

英文摘要

Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail to sufficiently exploit rich cross-frame global spatial information in videos. Even existing multi-level metric schemes only impose parallel prototype constraints on intermediate layers, without progressive supervision along the full feature pipeline, which results in limited generalization ability of the learned class prototypes. Inspired by this, we present a novel method, hierarchical metric learning for few-shot action recognition (HML-FSAR). First, a spatial-enhanced module is developed to capture cross-frame global spatial representations. Combined with temporal MHA, heterogeneous alignment, spatial-temporal feature fusion and dictionary learning modules, it constructs the complete feature processing pipeline. Second, a hierarchical metric learning (HML) strategy is embedded into HML-FSAR. Composed of center metric, alignment metric, contrastive metric, dictionary metric and prototype metric, HML imposes progressive multi-stage complementary constraints from frame-level representations to final class prototypes, so as to jointly optimize feature compactness, heterogeneous spatial-temporal alignment, inter-class discriminability and anti-noise robustness. The proposed HML-FSAR method is validated on five widely-used FSAR datasets, and experimental results fully demonstrate its effectiveness.

↑