通过混合交叉注意力网络的知识蒸馏和动态 INT8 量化实现高效视听事件识别
Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network
浏览论文内容
中文总结 AI 辅助
研究针对视听事件识别模型在边缘设备部署的难题,提出结合架构压缩、知识蒸馏和动态 INT8 量化的框架。构建教师与学生模型,经知识蒸馏训练学生模型,再用动态量化减小模型大小。实验表明该框架有效平衡准确率与效率,利于在资源受限平台部署。
中文摘要 AI 辅助
视听事件识别(AVER)通过基于变压器的多模态架构取得了显著的性能提升。然而,这些模型的高计算复杂度、大内存占用和推理成本阻碍了它们在边缘和资源受限设备上的部署。本文提出了一个高效的压缩框架,通过结合架构模型压缩、知识蒸馏和动态 INT8 量化,用于基于混合交叉注意力的视听事件识别。一个高容量的教师模型集成了用于视觉表示学习的 VideoMAE、用于音频特征提取的音频频谱变压器(AST)和用于多模态特征集成的混合交叉注意力融合网络。通过减少隐藏特征维度、注意力头数量和前馈网络大小,同时保留整体网络架构,构建了一个轻量级的学生模型。使用知识蒸馏训练学生模型,以有效地从教师模型中转移判别性知识。最后,应用动态 INT8 训练后量化进一步减小模型大小,以实现高效部署。在视听事件(AVE)数据集上的实验结果表明,所提出的框架将多模态融合模块中的可训练参数数量减少了 59.06%,与教师模型相比,分类准确率仅下降了 2.14%。此外,动态 INT8 量化将模型大小从 10.71MB 减少到 2.04MB,同时保持了有竞争力的识别性能。这些结果表明,所提出的框架在识别准确率和计算效率之间提供了有效的权衡,使其成为在资源受限的边缘 AI 平台上部署的有前途的解决方案。
英文摘要
Audio-visual event recognition (AVER) has achieved significant performance improvements through transformer-based multimodal architectures. However, the high computational complexity, large memory footprint, and inference cost of these models hinder their deployment on edge and resource-constrained devices. This paper presents an efficient compression framework for hybrid cross-attention-based audiovisual event recognition by combining architectural model compression, knowledge distillation, and dynamic INT8 quantization. A high-capacity teacher model integrates VideoMAE for visual representation learning, the Audio Spectrogram Transformer (AST) for audio feature extraction, and a hybrid cross-attention fusion network for multimodal feature integration. A lightweight student model is constructed by reducing the hidden feature dimension, the number of attention heads, and the feedforward network size while preserving the overall network architecture. The student model is trained using knowledge distillation to effectively transfer discriminative knowledge from the teacher. Finally, dynamic INT8 post-training quantization is applied to further reduce the model size for efficient deployment. Experimental results on the Audio-Visual Event (AVE) dataset show that the proposed framework reduces the number of trainable parameters in the multimodal fusion module by 59.06%, with only a 2.14% decrease in classification accuracy compared with the teacher model. Furthermore, dynamic INT8 quantization reduces the model size from 10.71 MB to 2.04 MB while maintaining competitive recognition performance. These results demonstrate that the proposed framework provides an effective trade-off between recognition accuracy and computational efficiency, making it a promising solution for deployment on resource-constrained edge AI platforms.
发表机构
- University of Porto, Porto, Portugal(葡萄牙波尔图大学)
- ResoSight, Montreal, Canada(ResoSight公司)
机构由 AI 辅助整理,请以论文原文为准。