arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03611cs.AIcs.MM

面向含不完整观测的多模态情感分析,重新思考模态可靠性

Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations

Chunlei Meng, Jacqueline J. Pang, Pengbin Feng, Zhenyu Yu, Chun Ouyang, Zhongxue Gan

AI总结:

针对含不完整观测的多模态情感分析,本文提出显式建模模态可靠性的MRCF框架,缓解可靠性不匹配与传播偏差,在多个公开情感数据集上取得优异性能。

AI中文摘要:

多模态情感分析(Multimodal Sentiment Analysis, MSA)整合文本、音频和视觉信息以推断人类情感,但现实中的多模态观测往往存在不完整情况。现有针对不完整观测的MSA方法主要遵循两种范式:基于重建的方法从观测到的模态中恢复缺失信息;联合表示方法直接从不完整输入中学习。尽管这些方法有效,但通常仅在表示学习或融合设计中隐式处理模态可靠性,而非显式建模。我们认为,在不完整观测场景下,模态可靠性是核心变量,未对其显式建模会引发两个相关问题:一是可靠性不匹配,即每个模态保留的情感证据随样本和缺失率变化;二是可靠性传播偏差,即退化模态的信息可能对跨模态交互和预测性能产生不利影响。为解决这些问题,我们提出MRCF(Modality Reliability-Calibrated Framework,面向含不完整观测的MSA的模态可靠性校准框架),其包含三个核心部分:可靠性感知分支,从模态内质量线索和跨模态语义一致性估计样本特定的模态可靠性;可靠性引导交互分支,利用估计的分数调节跨模态信息流;可靠性校准融合模块,整合可靠性与语义线索以生成最终预测。在CMU-MOSI、CMU-MOSEI和CH-SIMS数据集上的实验表明,MRCF在标准不完整观测协议下实现了优异性能;进一步分析提供证据显示,显式可靠性建模有助于缓解交互与融合过程中的可靠性不匹配和可靠性传播偏差。

英文摘要:

Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-based methods recover missing information from observed modalities, while joint-representation methods learn directly from incomplete inputs. Although effective, these methods usually treat modality reliability only implicitly within representation learning or fusion design rather than modeling it explicitly. We argue that modality reliability is a central variable in incomplete-observation settings. Failure to model it explicitly gives rise to two related issues. The first is reliability mismatch, in which the affective evidence retained by each modality varies across samples and missing rates. The second is reliability propagation bias, in which messages from degraded modalities may adversely affect cross-modal interaction and predictive performance. To address these issues, we propose MRCF, a Modality Reliability-Calibrated Framework for MSA with incomplete observations. MRCF contains a Reliability-Aware Branch that estimates sample-specific modality reliability from intramodal quality cues and cross-modal semantic consistency, a Reliability-Guided Interaction Branch that uses the estimated scores to modulate cross-modal information flow, and a Reliability-Calibrated Fusion Module that integrates reliability and semantic cues for final prediction. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show that MRCF achieves strong performance under standard incomplete-observation protocols. Further analyses provide evidence that explicit reliability modeling helps mitigate reliability mismatch and reliability propagation bias during interaction and fusion.

↑