arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向动作识别的基于自适应样本生成的通道级动态知识蒸馏

Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition

Ping Li, Chenhao Ping, Jie Song, Mingli Song

arXiv 2608.03100首次发表:更新:

AI 中文总结

针对现有动作识别知识蒸馏方法的样本固定与通道蒸馏强度统一问题,提出ASCD方法,通过自适应样本生成与通道级动态蒸馏实现最优特征对齐,在多数据集上取得SOTA性能。

AI 中文摘要

知识蒸馏(KD)为压缩大型动作识别模型提供了一条颇具前景但尚未充分探索的路径。然而,现有的KD方法存在两个关键局限:1)依赖固定输入样本,导致冻结的教师模型(更大的模型)与可学习的学生模型(更小的模型)之间的特征对齐效果欠佳;2)对所有通道采用统一的蒸馏强度,无法考虑它们在不同训练周期中捕捉不同知识(如运动节奏或幅度)的重要性差异。这促使我们开发了自适应样本感知通道级动态(ASCD)KD方法,该方法分为两个阶段运行。首先,我们使用自适应样本生成模块,通过结合样本梯度的语义来创建更新后的样本,样本梯度由最小化特征损失得出,该特征损失由各层的通道质心频率差异加权;同时,通过对频率特征应用高斯掩码来保留关键的运动相关细节。其次,我们采用通道级动态蒸馏模块,在这些生成的样本上训练学生模型,训练过程受样本梯度和特征频率的引导。为提高效率,样本是定期更新而非每个周期更新。在三个视频基准数据集(UCF101、Kinetics-400、Something-Something-v2)和两个图像数据集(CIFAR-100、ImageNet)上进行的大量实验表明,我们的方法达到了当前最优性能。代码可在此https URL获取。

英文摘要

Knowledge Distillation (KD) offers a promising yet underexplored path for compressing large action recognition models. However, existing KD methods suffer from two key limitations: 1) reliance on fixed input samples leads to suboptimal feature alignment between the frozen teacher (larger model) and the learnable student (smaller model), and 2) applying a uniform distillation strength for all channels fails to account for their varying importance in capturing distinct knowledge (e.g., motion tempo or magnitude) across training epochs. This motivates us to develop an Adaptive Sample-aware Channel-wise Dynamic (ASCD) KD approach, which operates in two stages. First, we use an adaptive sample generation module to create updated samples by incorporating semantics from sample gradients, which are derived by minimizing a feature loss weighted by channel centroid frequency differences at each layer. Meanwhile, crucial motion-related details are preserved by applying a Gaussian mask to frequency features. Second, we employ a channel-wise dynamic distillation module to train student on these generated samples, guided by sample gradients and feature frequencies. For efficiency, samples are updated periodically rather than per epoch. Extensive experiments on three video benchmarks (UCF101, Kinetics-400, Something-Something-v2) and two image datasets (CIFAR-100, ImageNet) demonstrate the state-of-the-art performance of our method. Code is available at https://github.com/mlvccn/ASCD_KD_Action.

CommentsAccepted in ACM MM2026, 16 pages, 7 figures

DOI:10.1145/3767308.3836050

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑