基于理由引导学习的多模态情感识别
Rationale-Guided Learning for Multimodal Emotion Recognition
浏览论文内容
中文总结 AI 辅助
针对多模态情感识别忽略人类因果推理的问题,提出理由引导学习框架,结合双加工理论与MLLM生成的理由训练模型,在IEMOCAP和MELD基准取得最优性能,且推理无额外开销。
中文摘要 AI 辅助
对话中的多模态情感识别(MERC)需要理解语言与非语言线索之间的复杂交互。然而,大多数现有方法从根本上将此问题视为直接的输入-输出(多模态线索-情感标签)映射问题,忽略了人类在解读情感时所使用的因果推理。我们提出理由引导学习(RGL),这是一个新颖的框架,将MERC转化为受认知启发的推理任务。基于双加工理论,我们将情感推理分解为三个方面:直觉(即时感知,系统1)、情境(情境分析,系统2)和整合(两者的综合)。我们利用多模态大语言模型(MLLM)离线生成结构化理由,这些理由被编码为记忆,通过将内部表示与类人推理模式对齐来指导模型训练。我们的最终模型在推理时无需任何MLLM开销。实验结果表明,RGL在IEMOCAP和MELD基准上实现了最先进的性能。此外,为了解释,我们证明该模型的内部特征能有效为未见测试样本检索语义正确的理由,验证了其理由推理能力。
英文摘要
Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches fundamentally treat this as a direct input-output (multimodal cues-emotion labels) mapping problem, overlooking the causal reasoning that humans use when interpreting emotions. We propose rationale-guided learning (RGL), a novel framework that transforms MERC into a cognitively-inspired reasoning task. Based on dual-process theory, we decompose emotional reasoning into three facets: Intuitive (immediate perception, System 1), Contextual (situational analysis, System 2), and Integrative (synthesis of both). We leverage an MLLM offline to generate structured rationales, which are encoded as memories to guide model training via aligning internal representations with human-like reasoning patterns. Our final model operates without any MLLM overheads at inference time. Experimental results show that RGL achieves state-of-the-art performance on the IEMOCAP and MELD benchmarks. Further, for interpretation, we demonstrate that the model's internal features effectively retrieve semantically correct rationales for unseen test samples, validating its rationale reasoning capabilities.