arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多模态情感分析的语义对齐结构抽象

Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

Wei Chen, Junkai Li, Tongguan Wang, Hui Liu, Feiyue Xue, Chuanxiang Ma, Ying Sha

arXiv 2607.27790首次发表:更新:

发表机构

Huazhong Agricultural University; Hubei University(华中农业大学; 湖北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有多模态情感分析方法无法建模情感语义的局限,提出SentiLLM框架,通过双流显著性-上下文校准机制实现语义对齐结构抽象,在四个数据集上取得优异性能。

AI 中文摘要

多模态情感分析(Multimodal Sentiment Analysis, MSA)旨在通过整合自然语言与非语言模态来解读复杂的人类情感。非语言模态与自然语言存在结构同构性,两者均可视为随时间演化的特征序列,这种同构性使得非语言模态可被转换为类文本标记以实现统一语义推理。专为理解和生成序列数据设计的大语言模型(Large Language Models, LLMs)因此可被用于解读复杂的情感序列。然而,现有的基于LLM的方法主要捕获低级表层特征,无法建模由结构变化和上下文交互产生的情感语义。为解决这一局限,我们提出SentiLLM,这是一个利用语义对齐结构抽象将连续原始信号提炼为紧凑、语义有意义标记的统一框架。具体而言,我们引入双流显著性-上下文校准机制,该机制将非语言特征序列解耦为焦点流和环境流:焦点流在文本先验引导下捕获显著的情感变化(如面部表情),环境流则表征稳定的背景状态。通过将这些动态情感变化与背景状态校准,SentiLLM可有效将非语言模态投影到统一语义空间,使其能被LLM自然理解。作为即插即用模块,SentiLLM仅需少量可训练参数即可显著提升判别性能,在MOSI、MOSEI、CH-SIMS和CH-SIMS v2四个数据集上取得了优异性能,证明了结构抽象范式在MSA中的有效性。我们的代码可在该网址获取:[此链接]

英文摘要

Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal modalities share a structural isomorphism with natural language, as both can be viewed as feature sequences evolving over time. This isomorphism enables the transformation of non-verbal modalities into text-like tokens for unified semantic reasoning. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized to interpret complex affective sequences. However, existing LLM-based methods primarily capture low-level superficial features, failing to model affective semantics arising from structural variations and contextual interactions. To address this limitation, we propose \textbf{SentiLLM}, a unified framework that leverages \textit{Semantic-Aligned Structural Abstraction} to distill continuous raw signals into compact, semantically meaningful tokens. Specifically, we introduce a \textit{Dual-Stream Salience-Context Calibration Mechanism}, which disentangles non-verbal feature sequences into a focus stream and an ambient stream. The focus stream captures salient sentiment shifts (e.g., facial expressions) guided by textual priors, while the ambient stream characterizes stable background states. Through calibrating these dynamic sentiment shifts against background states, SentiLLM effectively projects non-verbal modalities into a unified semantic space, making them naturally understandable for LLMs. Serving as a plug-and-play module, SentiLLM significantly improves discriminative performance with only a small number of trainable parameters. Our method achieves superior performance on four datasets, MOSI, MOSEI, CH-SIMS, and CH-SIMS v2, demonstrating the effectiveness of the structural abstraction paradigm in MSA. Our code is available at: \href{https://github.com/especiallyW/SentiLLM}.

CommentsAccepted by MM 2026

DOI:10.1145/3767308.3835866

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑