arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19826eess.AS

共识引导的共享-特定三视图学习用于语音情感识别

Consensus-Guided Shared-Specific Tri-View Learning for Speech Emotion Recognition

Bing Huang, Yujian Ma, Xikun Lu, Xianquan Jiang, Jinqiu Sang

AI总结:

针对语音情感识别中多视图融合冗余与互补信息丢失问题,提出TriCGF方法,通过共识学习与门控集成建模语谱图、MFCC和HuBERT,在IEMOCAP和EmoDB上取得最优性能。

AI中文摘要:

语音情感识别(SER)受益于异构声学表征,但源自同一话语的视图既包含重叠的情感证据,也包含依赖于表征的线索。直接融合因此可能传播冗余信息或掩盖互补细节。为解决此问题,我们提出三视图共识引导融合(TriCGF),用于联合建模语谱图、梅尔频率倒谱系数和HuBERT表征。TriCGF在融合前将每个视图组织为公共组件和视图特定组件。跨视图共识学习将公共组件聚合为全局参考,而视图门控集成自适应地将该参考与每个视图特定组件结合。软差异正则化进一步抑制过度的信息重叠。在说话人独立评估下,TriCGF在IEMOCAP上达到74.19%的加权准确率(WA)和75.17%的未加权准确率(UA),在EmoDB上达到94.36%的WA和94.28%的UA,在两个数据集上均优于代表性的SER方法。

英文摘要:

Speech emotion recognition (SER) benefits from heterogeneous acoustic representations, but views derived from the same utterance contain both overlapping emotional evidence and representation-dependent cues. Direct fusion may therefore propagate redundant information or obscure complementary details. To address this issue, we propose Tri-view Consensus-Guided Fusion (TriCGF) for jointly modeling spectrogram, Mel-frequency cepstral coefficients, and HuBERT representations. TriCGF organizes each view into common and view-specific components before fusion. Cross-view Consensus Learning aggregates the common components into a global reference, while View-wise Gated Integration adaptively combines this reference with each view-specific component. A soft difference regularizer further discourages excessive information overlap. Under speaker-independent evaluation, TriCGF achieves 74.19% weighted accuracy (WA) and 75.17% unweighted accuracy (UA) on IEMOCAP, and 94.36% WA and 94.28% UA on EmoDB, outperforming representative SER methods on both datasets.

补充信息

↑