发表机构
Beijing University of Posts and Telecommunications; Zhongguancun Academy; Tsinghua University(北京邮电大学; 中关村学院; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多模态情感分析中模态不完整的问题,提出MIDAS框架,通过互信息解缠与不确定性感知融合实现鲁棒的多模态表示,在多个数据集上取得优于基线的性能。
AI 中文摘要
大多数现有的多模态情感分析方法假设可以获取完整的多模态输入,但实际应用中经常遇到模态不完整或损坏的情况,这构成了关键挑战。尽管已提出多种方法来解决该问题,但它们主要依赖数据插补和启发式协调约束,无法有效从不完整多模态数据中提取和利用与任务相关的信息。为应对这一挑战,我们提出了一个名为Mutual Information Disentanglement with Uncertainty-Aware fuSion(MIDAS)的统一框架,该框架能在不完整条件下有效重构多模态表示。MIDAS采用变分建模策略,用多元高斯隐变量表示每个模态,并进一步将其分解为共享因子和专属因子。为获得可靠的表示,我们设计了一个极小极大目标,即最小化共享空间与专属空间之间的互信息以实现稳定解缠,同时最大化跨模态共享空间之间的互信息以增强语义对齐。此外,我们引入了不确定性感知融合机制,利用后验方差作为可靠性指标,在融合过程中自适应加权隐特征,确保即使在模态不完整时也能实现鲁棒集成。在三个广泛使用的数据集上进行的大量实验表明,MIDAS在各种不完整设置下均优于竞争性基线,取得了强劲且一致的性能提升,证明了其在不完整数据场景下的有效性和鲁棒性。
英文摘要
Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although several methods have been proposed to tackle this issue, they mainly rely on data imputation and heuristic coordination constraints, which fail to effectively extract and leverage task-relevant information from the incomplete multimodal data. To address this challenge, we propose a unified framework termed Mutual Information Disentanglement with uncertainty-Aware fuSion (MIDAS), which effectively restructures multimodal representations under incomplete conditions. MIDAS adopts a variational modeling strategy to represent each modality with multivariate Gaussian latent variables and further decomposes them into shared and exclusive factors. To obtain reliable representations, we design a minimax objective that minimizes the mutual information between shared and exclusive spaces for stable disentanglement, while maximizing the mutual information among shared spaces across modalities to enhance semantic alignment. In addition, an uncertainty-aware fusion mechanism is introduced, where posterior variance is leveraged as a reliability indicator to adaptively weight latent features during fusion, ensuring robust integration even when modalities are incomplete. Extensive experiments on three widely used datasets show that MIDAS achieves strong and consistent performance gains over competitive baselines across a wide range of incomplete settings, demonstrating its effectiveness and robustness for incomplete data scenarios.
CommentsAccepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026