SemMSA:基于潜在语义辅助的鲁棒多模态情感分析(数据不完整)
SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data
浏览论文内容
中文总结 AI 辅助
提出SemMSA框架,利用大语言模型构建潜在语义,通过跨模态语义精炼与谱对齐处理不完整多模态数据,在SIMS、MOSI和MOSEI基准上取得最先进性能。
中文摘要 AI 辅助
近年来,多模态情感分析(MSA)的研究聚焦于在数据不完整的情况下,从语言、视觉和声学模态中学习以推断人类情感。大多数研究通常通过重建模态特征或设计复杂的融合机制来补偿缺失信息。然而,由于在部分观测的多模态证据中缺乏高层语义基础,这些方法仍然存在虚假生成和噪声引导的问题。为解决这些问题,我们提出了SemMSA,一种潜在语义辅助框架,利用大语言模型(LLMs)构建丰富的情感相关语义,并通过无锚点谱对齐与所有模态完全集成。该框架主要由跨模态语义精炼(CSR)和跨模态谱对齐(CSA)组成。具体而言,CSR首先通过相应的适配器自适应地提取视觉和声学表示,在冻结的LLM嵌入空间中与语言形成统一的多模态前缀。然后,通过令牌高效的潜在精炼过程,迭代地产生连续的判别性语义状态,无需解码显式文本。接下来,CSA通过增强其核Gram矩阵的主谱分量,同时将精炼语义与所有模态对齐。这捕获了所有表示之间的全局非线性依赖关系,而无需依赖预定义的锚点模态。此外,实例级谱分离约束保留了跨样本判别性并缓解表示坍缩。在SIMS、MOSI和MOSEI基准上的大量实验表明,SemMSA达到了最先进的性能。
英文摘要
Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, these methods still suffer from spurious generation and noisy guidance due to the lack of high-level semantic grounding in partially observed multimodal evidence. To address these issues, we propose SemMSA, a latent semantic-aided framework that constructs rich sentiment-relevant semantics with LLMs, fully integrating with all modalities via anchor-free spectral alignment. It mainly consists of Cross-modal Semantic Refinement (CSR) and Cross-modal Spectral Alignment (CSA). Specifically, CSR first adaptively extracts visual and acoustic representations by corresponding adapters to form a unified multimodal prefix with language in the frozen LLM embedding space. It then iteratively produces continuous discriminative semantic states through a token-efficient latent refinement process without decoding explicit text. Next, CSA simultaneously aligns the refined semantics with all modalities by enhancing the dominant spectral component of their kernel Gram matrix. This captures global nonlinear dependencies among all representations without relying on a predefined anchor modality. In addition, an instance-level spectral separation constraint preserves cross-sample discriminability and mitigates representation collapse. Extensive experiments on SIMS, MOSI, and MOSEI benchmarks demonstrate that SemMSA achieves state-of-the-art performance.
发表机构
- Shandong University(山东大学)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。