发表机构
OPPO Research Institute(OPPO研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态情感识别中判别式融合丢失细节和生成式模型证据隐晦的问题,提出条件流框架BiCFlow-MER,实现生成式证据输运与冲突感知识别,在多个基准上超越现有方法。
AI 中文摘要
在多模态情感识别(MER)中,人类情感状态通过整合来自多种模态的互补线索进行推断。在音频-文本MER中,情感线索往往与说话者风格和词汇内容纠缠在一起,而跨模态分歧进一步使证据整合方式复杂化。在传统的判别式融合下,多模态证据被压缩为最终预测,模态特定线索和冲突信息保存不足。相比之下,在大型生成式情感模型中,情感推理通常嵌入语言解码中,使情感证据隐含且难以在结构化空间中验证。为解决这些局限,提出了BiCFlow-MER(双向条件流多模态情感识别),作为一种条件流框架,其中音频-文本MER被公式化为结构化情感空间内的生成式证据输运。在BiCFlow-MER中,面向情感的证据从说话者风格和词汇内容因素中解缠,以构建冲突感知的情感条件。在该条件引导下,每个话语通过双向校正流被输运到显式情感空间端点。候选情感通过输运端点的自适应原型云评分和与原始多模态条件的反向类别-条件一致性进行联合验证,实现冲突感知识别。BiCFlow-MER在IEMOCAP、MELD和零样本CASE基准上均优于所有对比方法。通过条件输运协调判别式识别和生成式证据建模,BiCFlow-MER定义了新的MER范式。
英文摘要
In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary cues from multiple modalities. In audio-text MER, affective cues are often entangled with speaker style and lexical content, while cross-modal disagreement further complicates how the evidence should be integrated. Under conventional discriminative fusion, multimodal evidence is compressed into a terminal prediction, with modality-specific cues and conflict information insufficiently preserved. In large generative affective models, by contrast, affective reasoning is typically embedded in language decoding, leaving emotion evidence implicit and difficult to verify in a structured space. To address these limitations, BiCFlow-MER (Bidirectional Conditional Flow for Multimodal Emotion Recognition) is proposed as a conditional-flow framework in which audio-text MER is formulated as generative evidence transport within a structured emotion space. Within BiCFlow-MER, emotion-oriented evidence is disentangled from speaker-style and lexical-content factors to construct a conflict-aware affective condition. Guided by this condition, each utterance is transported to an explicit emotion-space endpoint through a bidirectional rectified flow. Candidate emotions are jointly verified through adaptive prototype-cloud scoring of the transported endpoint and backward class-to-condition consistency with the original multimodal condition, enabling conflict-aware recognition. BiCFlow-MER is shown to outperform all compared methods across IEMOCAP, MELD, and the zero-shot CASE benchmark. By orchestrating discriminative recognition and generative evidence modeling through conditional transport, BiCFlow-MER defines a new MER paradigm.