arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20019cs.AI

针对未见模态组合的不完整多模态情感分析的对比混合提示学习

Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, Zixuan Yu, Wenxi Zhao, Yibei Liu, Qianle Zhang, Yangyang Wu, Mengying Zhu, Meng Xi

AI总结:

针对未见模态组合下的不完整多模态情感分析问题,提出CMPL模型,通过标签引导对比学习、带软路由的模态组合提示及三种提示对比策略,在三类数据集上较SOTA准确率提升超5%。

AI中文摘要:

近年来,不完整多模态情感分析受到广泛关注。现有方法通常假设数据随机缺失,或仅针对特定缺失模式设计,忽略了训练与测试阶段之间的模态组合不一致问题。然而在实际场景中,测试阶段常遇到训练阶段未出现过的模态组合,这导致模型泛化能力不足、性能不稳定。本文提出了**未见模态组合下的不完整多模态情感分析(IMSAUMC)**问题,旨在提升模型对未见模态组合的泛化能力。为解决该挑战,我们提出了针对IMSAUMC的**对比混合提示学习(CMPL)**模型。该模型引入标签引导的对比特征学习机制,以学习鲁棒且具有判别性的跨模态表示;此外,我们设计了带有软路由的模态组合提示,以促进对各类模态组合的更好学习;进一步,我们引入三种提示对比学习策略,使模型能有效学习对应未见模态组合的提示,从而显著增强模型在不同测试场景中的泛化能力。在三个广泛使用的数据集上开展的大量实验表明,与现有最优方法相比,CMPL的准确率提升超过5%。

英文摘要:

Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that were not present during the training phase, which leads to insufficient generalization capabilities and unstable performance. In this paper, we introduce the problem of Incomplete Multimodal Sentiment Analysis with Unseen Modality Combinations (IMSAUMC), aiming to enhance model generalization for unseen modality combinations. To address this challenge, we propose the model named $\textbf{C}$ontrastive $\textbf{M}$ixed $\textbf{P}$rompt $\textbf{L}$earning ($\textsf{CMPL}$) for IMSAUMC. It introduces a label-guided contrastive feature learning mechanism to learn robust and discriminative cross-modal representations. Additionally, we design modality-combination prompts with a soft router to facilitate better learning of various modality combinations. Furthermore, we introduce three prompt contrastive learning strategies, which enable effective learning of prompts corresponding to unseen modality combinations, thereby significantly strengthening the model's generalization capabilities in diverse testing scenarios. Extensive experiments on three widely used datasets demonstrate that $\textsf{CMPL}$ achieves more than a 5% improvement in accuracy compared to state-of-the-art approaches.

↑