arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29556cs.AI

跨模态情感理解:用于对话情感识别的Transformer-GAT方法

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

Jiaqi Qiao, Yifan Lyu, Xiujuan Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态情感识别中全局与局部上下文协同不足的问题,提出融合Transformer与图注意力网络的混合框架,在IEMOCAP和MELD数据集上取得领先性能。

中文摘要 AI 辅助

多模态情感识别是情感计算中的一个关键研究领域,在情感分析、智能客服和人机交互中具有广泛应用。然而,现有方法往往依赖单模态特征或简单的多模态融合,无法捕捉全局上下文与局部上下文之间的协同作用,这限制了模型性能和情感理解能力。为解决这一挑战,我们提出了Transformer-GAT,一种结合Transformer和图注意力网络的混合框架,以实现跨模态情感理解。Transformer用于捕捉全局语义信息,而图注意力网络用于建模模态之间的细粒度关系,从而增强情感特征的表示。在IEMOCAP和MELD数据集上的实验表明,我们的模型分别取得了72.45%和77.37%的加权F1分数,优于现有最先进的方法。这些结果证明,Transformer-GAT能够有效整合多模态特征,平衡全局与局部上下文,并提供更深入的情感洞察,为多模态情感计算提供了新方向。

英文摘要

Multimodal emotion recognition is a key research area in affective computing, with applications in sentiment analysis, intelligent customer service, and human-computer interaction. However, existing methods often rely on single-modal features or simple multimodal fusion, failing to capture the synergy between global and local contexts, which limits model performance and emotion understanding. To address this challenge, we propose Transformer-GAT, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding. The Transformer is used to capture global semantic information, while the Graph Attention Network is employed to model fine-grained relationships between modalities, thereby enhancing the representation of emotional features. Experiments on the IEMOCAP and MELD datasets show that our model achieves weighted F1 scores of 72.45% and 77.37%, outperforming state-of-the-art methods. These results demonstrate that Transformer-GAT effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion computing.

发表机构

  • Dalian University of Technology(大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑