arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用Whisper进行波斯语语音情感识别中的ASR适配与表示降维研究

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

Ali Shendabadi, Parnia Izadirad, Mostafa Salehi

arXiv 2608.05165首次发表:更新:

AI 中文总结

本研究针对低资源语言波斯语的语音情感识别问题,提出基于PCA降维的SER框架,实验表明该方法可提升性能并降低资源消耗,而ASR微调仅带来小幅增益。

AI 中文摘要

低资源语言的语音情感识别(SER)因标注数据有限仍是一个具有挑战性的问题。本研究探讨将Whisper用于波斯语SER,重点关注表示降维和特定语言的模型适配。我们提出了一个SER框架,其中从Whisper编码器提取的帧级嵌入使用主成分分析(PCA)进行降维,无需学习投影层,并大幅减少可训练参数数量;降维后的表示通过基于注意力的池化机制聚合,再由轻量级预测头进行分类。此外,我们研究在波斯语自动语音识别(ASR)任务上对Whisper进行微调是否能提升下游SER性能。在ShEMO数据集上采用说话人独立评估协议的实验表明,基于PCA的降维可持续提升情感识别性能,同时减少训练延迟和内存使用;ASR微调仅为SER带来小幅提升,表明在评估条件下,从语言适配到情感相关表示的迁移有限。这些发现为高效利用大型预训练语音模型开展低资源语言的情感识别提供了实用见解。

英文摘要

Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and language-specific model adaptation. We propose a SER framework in which frame-level embeddings extracted from the Whisper encoder are reduced in dimensionality using PCA, eliminating the need for learned projection layers and substantially reducing the number of trainable parameters. The reduced representations are aggregated using an attention-based pooling mechanism and classified with a lightweight prediction head. In addition, we investigate whether fine-tuning Whisper on a Persian automatic speech recognition (ASR) task improves downstream SER performance. Experiments conducted on the ShEMO dataset under a speaker-independent evaluation protocol show that PCA-based dimensionality reduction consistently improves emotion recognition performance while reducing training latency and memory usage. ASR fine-tuning yields only modest gains for SER, suggesting limited transfer from language adaptation to emotion-related representations under the evaluated conditions. These findings provide practical insights into the efficient use of large pretrained speech models for emotion recognition in low-resource languages.

Comments6 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑