arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向对话语音情感识别的双尺度状态空间建模与说话人感知动态条件随机场

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen

arXiv 2608.22399首次发表:更新:

发表机构

National Taiwan University of Science and Technology(台湾科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出DSSM-CRF架构,通过双尺度状态空间模型与说话人感知动态CRF分离对话情感识别的跨说话人影响与说话人内部演化,在IEMOCAP和MELD数据集上取得了优于对照的识别性能。

AI 中文摘要

对话语音情感识别需协调跨时间尺度的声学证据,涉及两类交互过程:跨说话人上下文影响与说话人内部情感演化。本文提出仅使用音频的架构DSSM-CRF,明确分离这两类过程。双向状态空间模型编码帧级与对话级的融合自监督语音表征,使每个话语表征捕获局部韵律与所有说话人的上下文信息。解码器将每个说话人的话语排序为独立的动态条件随机场链,同一说话人链中的连续话语构成转移对,其得分结合语料库级转移矩阵与从两个上下文化话语预测的残差。辅助目标监督每对是否发生情感变化,但不参与维特比推理。因此,对话者轮次会影响上下文情感得分,而不会被视为其他说话人情感轨迹的转移。DSSM-CRF在IEMOCAP数据集上取得75.81%未加权准确率(UA)和74.90%加权准确率(WA),在MELD数据集上取得54.72%加权准确率(WA)和49.31%加权F1值(WF1)。匹配对照实验表明,说话人分解与CRF建模可带来互补增益。

英文摘要

Conversational speech emotion recognition must reconcile acoustic evidence across temporal scales with two interaction processes: cross-speaker contextual influence and within-speaker emotion evolution. We propose DSSM-CRF, an audio-only architecture that explicitly separates these processes. Bidirectional state-space models encode fused self-supervised speech representations at frame and dialogue scales, so each utterance representation captures local prosody and context from all speakers. The decoder then orders each speaker's utterances into an independent dynamic conditional random field chain. Consecutive utterances in a speaker's chain form a transition pair whose score combines a corpus-level transition matrix with a residual predicted from the two contextualized utterances. An auxiliary objective supervises whether each pair changes emotion but does not participate in Viterbi inference. Thus, interlocutor turns affect contextual emotion scores without being treated as transitions in another speaker's emotion trajectory. DSSM-CRF achieves 75.81% UA and 74.90% WA on IEMOCAP, and 54.72% WA and 49.31% WF1 on MELD. Matched controls demonstrate complementary gains from speaker-wise factorization and CRF modeling.

Comments5 pages, 2 figures, 4 tables. Submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑