arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13024cs.CVcs.AI

标签引导的知识蒸馏用于3D-CNN动作识别

Label-Guided Knowledge Distillation for 3D-CNNs in Action Recognition

发表机构大数据分析国家工程实验室
查看机构详情
  • National Engineering Laboratory for Big Data Analytics (NEL-BDA)(大数据分析国家工程实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yanjiang Shi, Peng Zhao, Nan Qi, Guiqin Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对3D-CNN动作识别中特征蒸馏忽视时间维度的问题,提出标签引导知识蒸馏(LGKD),通过样本级与类别级蒸馏学习教师特征,在UCF101和HMDB51上取得竞争力结果。

中文摘要 AI 辅助

作为关键的模型压缩技术,知识蒸馏旨在将高容量教师模型的知识迁移到轻量级学生模型,以提升后者的性能。在本工作中,我们回顾了针对3D-CNN的特征知识蒸馏,并观察到视频分析中的大多数特征蒸馏方法仅是图像分析中方法的简单改编,往往忽略了视频特征在时间维度上的差异。为解决这一问题,我们提出了标签引导的知识蒸馏(Label-Guided Knowledge Distillation, LGKD),利用真实标签来指导学生模型特征的蒸馏。我们的方法包含两个组成部分:样本级蒸馏和类别级蒸馏,使学生模型能够在两个层面上学习教师模型的特征表示。样本级蒸馏利用标签信息和教师的概率分布,引导学习对时间准确性有显著影响的特征,同时减轻噪声。与此同时,类别级特征蒸馏采用原型网络(prototype network)进一步捕捉同一类别内样本间的关系知识,增强学生模型学习更高维语义信息的能力,并提升模型泛化性能。为证明我们方法的有效性和优越性,我们在两个基准动作识别数据集UCF101和HMDB51上进行了全面实验,取得了有竞争力的结果。

英文摘要

As a key model compression technique, knowledge distillation aims to transfer knowledge from a high-capacity teacher model to a lightweight student model for enhancing the latter's performance. In this work, we reviewed the feature knowledge distillation for 3D-CNNs and observed that most feature distillation methods in video analysis are simple adaptations of those used in image analysis, often neglecting the differences of video features in the temporal dimension. To address this issue, we proposed Label-Guided Knowledge Distillation (LGKD) to guide the distillation of student model features using ground truth labels. Our method entails two components: sample-wise distillation and class-wise distillation, enabling the student model to learn feature representation of the teacher model at two levels. Sample-wise distillation utilizes label information and the teacher's probability distribution to guide the learning of features that significantly impact temporal accuracy while mitigating noise. Meanwhile, class-wise feature distillation employs a prototype network to further capture the relational knowledge among samples within the same category, enhancing the student's ability to learn higher-dimensional semantic information and improving model generalization. To demonstrate the effectiveness and superiority of our method, we conducted comprehensive experiments on two benchmark action recognition datasets, UCF101 and HMDB51, achieving competitive results.

补充信息

↑