arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于语音情感识别的可解释轻量级紧凑深度模型

Explainable Lightweight Compact Deep Models for Speech Emotion Recognition

Nelly Elsayed

arXiv 2607.16803首次发表:更新:

AI 中文总结

针对语音情感识别中现有方法存在的问题,提出基于紧凑卷积神经网络架构的可解释轻量级框架,利用对数梅尔频谱图和注意力统计池化,结合基于梯度的类激活映射,在SAVEE数据集上实验,平衡了识别精度、计算效率和模型透明度。

AI 中文摘要

语音情感识别(SER)是众多以人类为中心的应用中的重要组成部分。在医疗和决策支持环境中,人们越来越关注不仅能实现准确情感识别,还能支持透明预测和高效部署的模型。然而,现有许多SER方法依赖复杂深度学习架构,限制了可解释性并增加计算成本。本文提出基于紧凑卷积神经网络架构的可解释轻量级语音情感识别框架。利用对数梅尔频谱图表示捕捉语音特征,采用注意力统计池化强调情感显著时间段,通过基于梯度的类激活映射提高模型透明度。在SAVEE情感语音数据集上的实验表明,该框架在保持紧凑架构且参数比现有许多SER模型少得多的情况下,实现了有竞争力的识别性能。结果表明,高效卷积架构与可解释分析相结合,能在识别精度、计算效率和模型透明度之间实现实际平衡。

英文摘要

Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer service, and human-omputer interaction. In medical and decision-support settings, there is increasing interest in models that not only achieve accurate emotion recognition but also support transparent predictions and efficient deployment. However, many existing SER approaches rely on complex deep learning architectures that limit interpretability and increase computational cost. This paper presents an explainable and lightweight speech emotion recognition framework based on a compact convolutional neural network architecture. The proposed approach utilizes log-Mel spectrogram representations to capture spectro-temporal speech characteristics and employs attentive statistics pooling to emphasize emotionally salient temporal segments. To improve model transparency, gradient-based class activation mapping (Grad-CAM) is incorporated to visualize the time-frequency regions that influence the model's predictions. Experimental evaluation on the SAVEE emotional speech dataset demonstrates that the proposed framework achieves competitive recognition performance while maintaining a compact architecture with significantly fewer parameters than many existing SER models. The results indicate that efficient convolutional architectures combined with interpretable analysis can provide a practical balance between recognition accuracy, computational efficiency, and model transparency.

CommentsAccepted in the IEEE ICMLA 2026 Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑