发表机构
College of Computer Science and Technology, Jilin University; Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(吉林大学计算机科学与技术学院; 吉林大学符号计算与知识工程教育部重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出结合多特征编码与注意力融合的多模态情感识别框架,在MELD、IEMOCAP数据集上性能优于基线,可有效捕捉语音与面部情感线索,实用且可泛化。
AI 中文摘要
多模态情感识别因在人机交互、远程教育及医疗保健领域的重要性而受到日益关注。本文提出一种新型多模态情感识别框架,该框架将丰富的音频与视觉特征提取与基于注意力的融合策略相结合。对于音频,我们提取三类互补特征:来自Wav2Vec2的语义嵌入、MFCC特征,以及音高、能量、节奏等统计声学描述符,这些特征经对齐后通过BiLSTM融合以捕捉时序依赖关系。对于视频,我们提出ResNet50-BiLSTM架构,结合深度残差学习与序列建模,从面部序列中提取具表现力的时空特征。为增强多模态协同,我们引入基于多头注意力的特征级融合机制,使模型能自适应权衡各模态的贡献。在MELD与IEMOCAP数据集上开展的实验表明,我们的模型在准确率与鲁棒性上均显著优于基线模型;此外, ablation研究显示,基于注意力的融合策略在不平衡数据设置下可显著提升性能。我们的研究结果表明,所提框架能有效捕捉语音与面部表情中的多样情感线索,为现实世界多模态情感识别任务提供了实用且可泛化的方法。
英文摘要
Multimodal emotion recognition has attracted growing interest due to its importance in human-computer interaction, remote education, and healthcare. This paper proposes a novel multimodal emotion recognition framework that integrates rich audio and visual feature extraction with an attention-based fusion strategy. For audio, we extract three complementary feature types: semantic embeddings from Wav2Vec2, MFCC features, and statistical acoustic descriptors such as pitch, energy, and rhythm. These are aligned and fused via a BiLSTM to capture temporal dependencies. For video, we propose a ResNet50-BiLSTM architecture that combines deep residual learning and sequential modeling to extract expressive spatiotemporal features from facial sequences. To enhance multimodal synergy, we introduce a feature-level fusion mechanism based on multi-head attention, allowing the model to adaptively weigh contributions across modalities. Experiments conducted on the MELD and IEMOCAP datasets demonstrate that our model significantly outperforms baselines in both accuracy and robustness. Furthermore, ablation studies show that the attention-based fusion strategy significantly improves performance in unbalanced data settings. Our findings suggest that the proposed framework effectively captures diverse emotional cues from speech and visual expressions, and offers a practical and generalizable approach for real-world multimodal emotion recognition tasks.
Comments15 pages, 6 figures, 6 tables. Pre-peer-review version. The final published version appears in ICONIP 2025, Lecture Notes in Computer Science, vol. 16312, pp. 142-157 (2026)
Journal refNeural Information Processing (ICONIP 2025), Lecture Notes in Computer Science, vol. 16312, pp. 142-157. Springer, Singapore (2026)
DOI:10.1007/978-981-95-4384-7_11