发表机构
Central South University of Forestry and Technology; State University of New York, New Paltz(中南林业科技大学; 纽约州立大学新帕尔茨分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态情感识别中模态缺失导致性能下降的问题,本文提出C²MOE框架,通过分解多模态知识为一致性与互补性分量、引入双分支预测机制及可学习重加权模块,在多个基准上实现了优于现有方法的性能。
AI 中文摘要
多模态对话情感识别(Multimodal Emotion Recognition in Conversations, MERC)的近期进展凸显其对完整多模态输入的依赖,但现实世界数据常因传输错误或用户行为导致模态缺失,严重降低模型性能。现有方法通过跨模态一致性学习提升鲁棒性,但大多忽略模态互补性,导致重建结果存在偏差。为解决该局限,本文提出C²MOE,一种面向不完整多模态情感学习的新型一致性与互补性引导的混合专家(Mixture of Experts)框架。该方法在原则性信息论框架内统一表征学习与缺失模态补全,具体而言,通过感知交互的专家将多模态知识分解为一致性与互补性分量:一致性通过最大化跨模态可预测性捕获,互补性通过最大化模态间条件熵保留。基于该分解,C²MOE引入双分支预测机制以在模态缺失时实现鲁棒补全:一致性分支通过最小化不确定性使补全特征与联合分布对齐,互补性分支通过熵最大化挖掘模态独有关键信息。最后,C²MOE采用可学习重加权模块,动态分配各专家输出的重要性分数,实现补全任务的鲁棒自适应融合。在多个MERC基准上的大量实验表明,C²MOE在各类缺失模态设置下均持续超越现有最优方法,验证了其鲁棒性与泛化性。
英文摘要
Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrading model performance. Existing methods enhance robustness via cross-modal consistency learning but largely ignore modality complementarity, leading to biased reconstructions. To address this limitation, we propose C$2$MOE, a novel Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion learning. Our approach unifies representation learning and missing modality imputation within a principled information-theoretic framework. Specifically, multimodal knowledge is factorized into consistency and complementarity components via interaction-aware experts. Consistency is captured by maximizing cross-modal predictability, while complementarity is preserved by maximizing conditional entropy between modalities. Building upon this decomposition, C$2$MOE introduces a dual-branch prediction mechanism for robust imputation under missing modalities. The consistency branch aligns imputed features with the joint distribution by minimizing uncertainty, and the complementarity branch exploits modality-unique cues via entropy maximization. Finally, C$2$MOE employs a learnable reweighting module that dynamically assigns importance scores to each expert's output, yielding a robust and adaptive fusion for imputation. Extensive experiments on multiple MERC benchmarks demonstrate that C$2$MOE consistently surpasses state-of-the-art methods across various missing-modality settings, validating its robustness and generalization.
Comments10 pages