XEmoGPT:一种具有线索级感知与推理的可解释多模态情感识别框架
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
- School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
- th Research Institute, China Electronics Technology Group Corporation(中国电子科技集团第五十四研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
XEmoGPT通过专门模块和大规模数据集提升多模态情感识别的可解释性和线索级推理能力。
AI中文摘要:
可解释的多模态情感识别在人机交互和社会媒体分析等应用中起着关键作用。然而,当前方法在线索级感知和推理方面面临两大挑战:1)通用模态编码器是为捕捉全局结构和通用语义进行预训练的,而非精细情感线索,导致对情感信号的敏感性有限;2)现有数据集通常在标注质量和规模之间存在权衡,导致对情感线索的监督不足,最终限制了线索级推理。此外,现有评估指标不足以评估线索级推理性能。为解决这些挑战,我们提出了eXplainable Emotion GPT(XEmoGPT),一种新的EMER框架,能够同时感知和推理情感线索。它结合了两个专门模块:视频情感线索桥(VECB)和音频情感线索桥(AECB),通过精心设计的任务增强视频和音频编码器以实现细粒度的情感线索感知。为了进一步支持线索级推理,我们构建了一个大规模数据集EmoCue,旨在教会XEmoGPT如何推理多模态情感线索。此外,我们引入了EmoCue-360,一个自动指标,通过语义相似性提取和匹配情感线索,并发布了EmoCue-Eval,一个包含400个专家标注样本的基准测试,涵盖多样化的心理场景。实验结果表明,XEmoGPT在情感线索感知和推理方面均表现出色。
英文摘要:
Explainable Multimodal Emotion Recognition plays a crucial role in applications such as human-computer interaction and social media analytics. However, current approaches struggle with cue-level perception and reasoning due to two main challenges: 1) general-purpose modality encoders are pretrained to capture global structures and general semantics rather than fine-grained emotional cues, resulting in limited sensitivity to emotional signals; and 2) available datasets usually involve a trade-off between annotation quality and scale, which leads to insufficient supervision for emotional cues and ultimately limits cue-level reasoning. Moreover, existing evaluation metrics are inadequate for assessing cue-level reasoning performance. To address these challenges, we propose eXplainable Emotion GPT (XEmoGPT), a novel EMER framework capable of both perceiving and reasoning over emotional cues. It incorporates two specialized modules: the Video Emotional Cue Bridge (VECB) and the Audio Emotional Cue Bridge (AECB), which enhance the video and audio encoders through carefully designed tasks for fine-grained emotional cue perception. To further support cue-level reasoning, we construct a large-scale dataset, EmoCue, designed to teach XEmoGPT how to reason over multimodal emotional cues. In addition, we introduce EmoCue-360, an automated metric that extracts and matches emotional cues using semantic similarity, and release EmoCue-Eval, a benchmark of 400 expert-annotated samples covering diverse emotional scenarios. Experimental results show that XEmoGPT achieves strong performance in both emotional cue perception and reasoning.