发表机构
University of Hamburg(汉堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对对话情感识别中现有方法的局限,提出SCoPE模块,利用说话者情感历史建模先验,纳入情感转移预测并采用转移感知融合机制,在多模态IEMOCAP数据集上取得优于现有模型的性能。
AI 中文摘要
在对话中,人类情感是短暂的,但往往会在多个话语中持续。例如,很少会在如快乐和愤怒等对比情绪之间瞬间切换,情绪倾向于平稳演变且具有说话者特异性。现有对话情感识别(ERC)方法多依赖显性证据,未充分建模非显性因素,在多模态场景下信号有噪声时模型易脆弱。为解决这些局限,引入说话者情感条件先验(SCoPE),它利用说话者情感历史并明确建模其先验用于后续情感分类。还纳入情感转移预测来指导模型平衡SCoPE先验和多模态证据,提出转移感知融合机制进行精确加权逻辑整合。实验结果表明该模型在多模态IEMOCAP数据集上性能优于近期最先进模型。
英文摘要
In conversations, human emotions are transient; however, they tend to persist across multiple utterances. For example, we rarely switch instantly between contrasting emotions such as happiness and anger. Instead, emotions tend to evolve smoothly, and these patterns are often speaker-specific. Some people might escalate, while others gradually cool down over time. Furthermore, when emotions change during a conversation, they are often driven by contextual factors, such as newly received information or unexpected events. Even though progress has been made in Emotion Recognition in Conversations (ERC), most existing approaches still rely heavily on overt evidence and do not sufficiently model these non-apparent factors. Especially in multimodal settings, this makes these models fragile when the signals are noisy (e.g., occluded faces, slang expressions, or microphone noise). To address these limitations, we introduce Speaker-Conditioned Priors over Emotions (SCoPE). SCoPE is a light weight module that utilizes the emotional history of each speaker and explicitly models their priors for use in subsequent emotion classification. Second, we incorporate emotion shift prediction, a well-established concept in ERC, to guide the model in balancing the priors from SCoPE and multimodal evidence. Finally, we propose a shift-aware fusion mechanism that performs precision-weighted logit integration between multimodal evidence and the speaker prior, forming a Bayesian-inspired product-of-experts formulation. This dynamic fusion allows the model to rely on historical priors when emotions persist and to prioritize multimodal evidence when shifts are likely. Experimental results show our model achieves superior performance over recent state-of-the-art models on the IEMOCAP dataset in multimodal settings.
CommentsUnder review at Cognitive Computation