发表机构
Jiangsu University; Griffith University; National University of Singapore; University of New South Wales; Fujian University of Technology; Xi’an Jiaotong-Liverpool University; German Research Center for Artificial Intelligence; Tsinghua University; Technical University of Munich; Imperial College London; Queen Mary University of London(江苏大学; 格里菲斯大学; 新加坡国立大学; 新南威尔士大学; 福建理工大学; 西交利物浦大学; 德国人工智能研究中心; 清华大学; 慕尼黑工业大学; 伦敦帝国学院; 伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对不完整多模态情感识别中现有方法忽略模态异质性的问题,提出PriMD框架,通过解耦语义与原语构建记忆库,在多数据集上实现SOTA性能与更强鲁棒性。
AI 中文摘要
多模态情感识别(MER)系统在实际场景中常面临模态缺失问题。现有方法通常将缺失模态作为整体进行生成、对齐或蒸馏,却忽略了各模态所承载信息的异质性,这种整体处理方式将可推断的共享语义与不确定的模态特定细节混合,导致表示不稳定并降低鲁棒性。为解决该问题,本文提出原语记忆蒸馏(PriMD)框架。与现有方法不同,PriMD从模态内视角出发,关注模态内不同类型信息在可恢复性上的差异。PriMD首先将跨模态共享语义与模态特定表示解耦,再将后者离散为可学习的语义原语以构建模态特定记忆库。当模态缺失时,PriMD采用师生框架:学生模型利用可用模态的共享语义作为查询动态检索原语,在受限记忆空间内补偿缺失的模态特定信息,并与教师模型对齐。在IEMOCAP、CMU-MOSI和CMU-MOSEI上的大量实验表明,PriMD在多种模态缺失设置下均达到了最先进性能,且鲁棒性持续更强,同时缓解了整体特征推断带来的不稳定性。我们的代码和项目网站分别位于this https URL和this https URL。
英文摘要
Multimodal Emotion Recognition (MER) systems often suffer from missing modalities in real-world scenarios. Existing methods usually generate, align, or distill missing modalities as a whole, overlooking the heterogeneous nature of the information carried by each modality. Such holistic treatment mixes inferable shared semantics with uncertain modality-specific details, yielding unstable representations and degrading robustness. To address this issue, we propose the Primitive Memory Distillation (PriMD) framework. Unlike existing methods, PriMD takes an intra-modal perspective and focuses on how different types of information within a modality differ in recoverability within each modality. PriMD first disentangles cross-modal shared semantics from modality-specific representations, and then discretizes the latter into learnable semantic primitives to construct modality-specific memory banks. When modalities are missing, PriMD is a teacher-student framework that the student model uses the shared semantics of available modalities as queries to dynamically retrieve primitives. It compensates for missing modality-specific information within a constrained memory space and aligns with the teacher model. Extensive experiments on IEMOCAP, CMU-MOSI, and CMU-MOSEI demonstrate that PriMD achieves state-of-the-art performance and consistently stronger robustness across a wide range of missing-modality settings, while mitigating the instability caused by holistic feature inference. Our code and project website are available at https://github.com/JiaqiZhang-Sengoku/PriMD and https://jiaqizhang-sengoku.github.io/PriMD/, respectively.
Comments19 Pages, 8 Figures, 13 Tables. Accepted to EMNLP 2026 Findings