SISER:基于熵对抗训练的说话人无关语音情感识别
SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training
浏览论文内容
中文总结 AI 辅助
本文针对语音情感识别的标注数据稀缺与说话人差异性问题,提出基于熵对抗训练的SISER方法,采用wav2vec 2.0与ECAPA-TDNN,在IEMOCAP数据集上取得优于基线及无说话人抑制的wav2vec 2.0的性能。
中文摘要 AI 辅助
语音情感识别(SER)面临两大核心挑战:标注数据稀缺与说话人差异性,二者均会阻碍情感识别系统的泛化能力。现有对抗方法虽能处理说话人差异性,但未能充分利用强大的预训练表示。本文提出SISER(说话人无关语音情感识别),将wav2vec 2.0作为特征编码器、ECAPA-TDNN作为说话人判别器,整合至基于熵的对抗训练方案中。wav2vec 2.0提供丰富的自监督表示,缓解对大量标注数据集的依赖;ECAPA-TDNN相比浅层分类器,能通过更强的对抗信号抑制说话人身份。在IEMOCAP数据集上评估,SISER的未加权准确率(UA)达60.63%,优于基线(51.15%)和无说话人抑制的wav2vec 2.0(56.46%), ablation实验表明,说话人分类器架构的选择是关键因素。
英文摘要
Speech emotion recognition (SER) faces two fundamental challenges: scarcity of labeled data and inter-speaker variability, both of which hinder generalization of emotion recognition systems. While prior adversarial approaches address speaker variability, they fall short in leveraging powerful pre-trained representations. We propose SISER (Speaker-Invariant Speech Emotion Recognition), integrating wav2vec 2.0 as a feature encoder and ECAPA-TDNN as a speaker discriminator within an entropy-based adversarial training scheme. wav2vec 2.0 provides rich self-supervised representations that alleviate dependency on large labeled datasets, while ECAPA-TDNN enables suppression of speaker identity via a stronger adversarial signal than shallow classifiers. Evaluated on IEMOCAP, SISER achieves a UA of 60.63%, outperforming the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%), with ablation emphasizing that the choice of speaker classifier architecture is a key factor.