arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16962cs.AI

情感原型引导融合用于开放词汇不完整多模态情感识别

Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition

Yichi Zhang, Shenyue Wang, Jing Luo, Chunyang Yu, Xinyu Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对开放词汇多模态情感识别中模态缺失问题,提出情感原型条件融合框架,通过构建情感原型库动态约束融合,显著优于现有方法。

中文摘要 AI 辅助

开放词汇多模态情感识别(OV-MER)旨在从多模态情感线索中生成开放的自然语言情感标签。然而,在现实场景中,由于采集设备的限制和用户隐私约束,难以获得完整且同步的模态数据。现有的OV-MER方法主要针对全模态输入设计,在模态缺失条件下无法进行有效的特征融合。同时,当前针对不完整模态设计的融合方法主要聚焦于固定标签识别场景,无法满足在OV-MER场景中由任意情感语义引导的情感线索融合需求。为应对这些挑战,本文提出了一种情感原型条件融合(APCF)框架,用于不完整开放词汇情感识别。作为一种无候选的生成式框架,APCF将模态贡献学习扩展到由任意情感语义引导的场景中。具体而言,我们构建了一个情感原型库,以显式建模不同情感对应的多模态贡献特征,从而在不同情感语义视角下为模态融合提供动态约束。基于可用模态特征进行条件检索和特征聚合。随后,将精炼后的融合情感表示输入到LLM解码器中,以生成开放词汇的情感标签。在OV-MERD+和MER-FG数据集上的实验表明,APCF显著优于最先进的基线方法。

英文摘要

Open-vocabulary multimodal emotion recognition (OV-MER) aims to generate open natural-language emotion labels from multimodal affective cues. In real-world scenarios, however, complete and synchronized modal data are difficult to obtain due to limitations of acquisition devices and user privacy constraints. Existing OV-MER methods are largely designed for full-modal inputs, and fail to perform effective feature fusion under modal missing conditions. Meanwhile, current fusion approaches designed for incomplete modalities mainly focus on fixed-label recognition context, and cannot satisfy the demand for fuse emotional cues guided with arbitrary emotion semantics in OV-MER context. To tackle these challenges, this paper proposes an Affect-Prototype-Conditioned Fusion (APCF) framework for incomplete open-vocabulary emotion recognition. As a candidate-free generative framework, APCF extends modal contribution learning to scenarios guided by arbitrary emotional semantics. Specifically, we construct an affect-prototype library to explicitly model multimodal contribution characteristics corresponding to diverse emotions, which provides dynamic constraints for modal fusion under different emotional semantic perspectives. Conditional retrieval and feature aggregation are conducted based on available modal features. The refined fused affective representations are then fed into an LLM decoder to produce open-vocabulary emotion labels. Experiments on the OV-MERD+ and MER-FG datasets demonstrate that APCF substantially outperforms state-of-the-art baselines.

发表机构

  • OPPO Research Institute(OPPO研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑