arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

情感作为分布:面向说话人无关多模态情感识别的联合效价-唤醒概率学习

Emotion as a Distribution: Joint Valence-Arousal Probability Learning for Speaker-Independent Multimodal Emotion Recognition

Tingyi Lin, Wen-Ren Yang, Kuanwei Chen

arXiv 2609.05755首次发表:更新:

发表机构

National Changhua University of Education; National Central University(国立彰化师范大学; 国立中央大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态情感识别将情感压缩为硬标签的问题,提出输出9×9效价-唤醒概率分布的双头系统,在IEMOCAP上超越Transformer基线3.0个UA点,并恢复环状模型结构。

AI 中文摘要

人类情感是分级且经常混合的,然而大多数多模态识别器将其压缩为单一硬标签。我们认为识别器应转而输出情感空间上的分布。我们的文本+语音系统,在分类决策之外,还输出一个$9\times9$的概率矩阵覆盖效价-唤醒平面,该矩阵在Kullback-Leibler/交叉熵目标下使用二维高斯软目标进行训练,旨在用于咨询支持。评估严格:采用说话人无关的5折留一会话交叉验证IEMOCAP,并进行轮转会话内部验证,主要指标仅在留出会话上计算。在固定的编码器-融合-头部流水线中,我们在匹配深度和宽度下比较了Transformer和状态空间(Mamba-1/2/3)骨干网络,在两个操作点($T\approx550$,$T\approx2750$)进行。所提出的双头系统在三个种子上达到73.0% $\pm$ 0.3的未加权准确率(单独重跑:72.1%),超过Transformer融合基线3.0个UA点(95%会话自助法置信区间[1.0,4.7];在配对t检验和会话级自助法下显著),在这些长度下没有延迟或内存优势;将约1M可训练前端替换为冻结的WavLM-Large特征(可学习层权重)将相同架构提升至76.6% $\pm$ 1.3。预先指定的对照诚实界定了声明范围:更简单的效价-唤醒辅助任务在噪声范围内复现了分类提升,而专门的回归头能略微更好地跟踪连续评分,因此该头部的特定价值在于归一化的情感分布本身。该分布恢复了环状模型:其质心跟踪效价和唤醒(CCC 0.66/0.66;主要为类间结构,类内跟踪较弱),其熵与分类评分者的模糊性弱但一致相关,而非维度离散度。

英文摘要

Human emotion is graded and frequently mixed, yet most multimodal recognizers collapse it onto a single hard label. We argue the recognizer should instead expose a distribution over affective space. Our text+speech system, alongside its categorical decision, emits a $9\times9$ probability matrix over the Valence-Arousal plane, trained with a two-dimensional Gaussian soft target under a Kullback-Leibler/cross-entropy objective, aimed at counseling support. Evaluation is strict: speaker-independent 5-fold leave-one-session-out IEMOCAP with rotating-session inner validation, headline metrics only on the held-out session. Within one fixed encoder-fusion-head pipeline we compare Transformer and state-space (Mamba-1/2/3) backbones at matched depth and width, at two operating points ($T\approx550$, $T\approx2750$). The featured dual-head system reaches 73.0% $\pm$ 0.3 unweighted accuracy over three seeds (separate rerun: 72.1%), exceeding the Transformer fusion baseline by 3.0 UA points (95% session-bootstrap CI [1.0,4.7]; significant under paired t-test and session-level bootstrap), with no latency or memory advantage at these lengths; swapping the ~1M trainable front-end for frozen WavLM-Large features (learnable layer weights) lifts the same architecture to 76.6% $\pm$ 1.3. Pre-specified controls scope the claims honestly: simpler valence-arousal auxiliaries reproduce the classification lift within noise, and a dedicated regression head tracks the continuous ratings slightly better, so the head's specific value is the normalized affect distribution itself. That distribution recovers the circumplex: its center of mass tracks valence and arousal (CCC 0.66/0.66; predominantly between-class structure, weaker within-class tracking), and its entropy is weakly but consistently linked to categorical rater ambiguity, not dimensional spread.

Comments18 pages, 8 figures. Submitted to Speech Communication. Code: https://github.com/brian10420/EchoMind-Mamba-VA-SER (archived: doi:10.5281/zenodo.21904810)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑