arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09924cs.CV

对话中的多模态情绪识别:基于类别自适应模态融合与情感几何

Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry

Oriol Marín, Roger Marí, Gloria Haro, Rafael Redondo

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种结合外观与几何视觉表示、类别自适应模态融合及效价-唤醒先验的ERC方法,在MELD和IEMOCAP上显著提升性能,验证了结构化面部线索与情感几何的互补价值。

中文摘要 AI 辅助

对话情绪识别(ERC)需要整合异构的文本、音频和视觉线索,同时考虑对话上下文和情绪动态。我们扩展了用于ERC的自蒸馏Transformer架构,采用外观+几何视觉表示、类别自适应模态融合以及用于情感转换的效价-唤醒先验。在MELD和IEMOCAP数据集上,与仅使用外观特征相比,几何增强的视觉表示分别将加权F1提高了0.27和4.36个百分点,而类别自适应融合在原始softmax门控基础上进一步提升了0.17和0.25个百分点。效价-唤醒先验在情感偏移话语上带来了0.30和0.74个准确率点的针对性改进,同时保持了稳定轮次的性能。这些结果表明,结构化面部线索、情绪依赖的模态权重和情感几何为多模态ERC提供了互补的益处。

英文摘要

Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transformer architecture for ERC with appearance+geometry visual representations, class-wise adaptive modality fusion, and a valence-arousal prior for affective transitions. On the MELD and IEMOCAP datasets, geometry-enhanced visual representations improve weighted F1 by 0.27 and 4.36 points over appearance-only features, respectively, while class-wise adaptive fusion provides further gains of 0.17 and 0.25 points over the original softmax gate. The valence-arousal prior yields targeted improvements of 0.30 and 0.74 accuracy points on emotionally shifted utterances while preserving performance on stable turns. These results indicate that structured facial cues, emotion-dependent modality weighting, and affective geometry provide complementary benefits for multimodal ERC.

发表机构

  • Eurecat, Centre Tecnològic de Catalunya(加泰罗尼亚技术中心Eurecat)
  • Universitat Pompeu Fabra(庞培法布拉大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑