发表机构
Fuzhou University; Harbin Institute of Technology (Shenzhen); Jiangxia University(福州大学; 哈尔滨工业大学(深圳); 江夏大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出MGSI多粒度情感集成框架,通过多尺度编码、文本引导对齐等优化,提升基于LLM的多模态情感分析性能,在四个公开基准上效果优于冻结LLM基线。
AI 中文摘要
多模态情感分析(MSA)旨在从文本、音频、视觉等异构输入中预测情感极性与强度。尽管大语言模型(LLM)为MSA提供了强大的语义先验,但如何有效融合音频和视觉信号仍是挑战。一个关键难点在于,音频和视觉情感线索在不同时间尺度上演化,而许多基于LLM的方法在将这些信号与文本融合前,通过浅层投影或粗池化压缩,这会削弱跨模态对齐并抹去细粒度情感信息。我们提出MGSI,一种面向基于LLM的MSA的多粒度情感集成框架。MGSI首先在短、中、长三个时间尺度上编码音频和视觉流,保留局部变化与全局情感趋势;随后通过文本引导对齐优化非文本特征,并应用感知极性和强度的增强,以更好处理模糊和近中性样本;最终将生成的多模态表示压缩为少量伪令牌,用于高效调控冻结的LLM。在四个公开基准上的实验表明,MGSI显著优于冻结LLM基线,且与强大的多模态方法相比仍具竞争力;进一步的消融和敏感性分析验证了多粒度时间建模、文本引导优化以及自适应情感校准的有效性。
英文摘要
Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision. While large language models (LLMs) offer strong semantic priors for MSA, effectively incorporating audio and visual signals effectively remains challenging. A key challenge is that audio and visual sentiment cues evolve over different temporal scales, yet many LLM-based methods compress these signals through shallow projection or coarse pooling before fusing them with text, which can weaken cross-modal alignment and erase fine-grained affective information. We propose MGSI, a multi-granularity sentiment integration framework for LLM-based MSA. MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends. It then refines non-text features through text-guided alignment, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples. The resulting multimodal representation is finally compressed into a small set of pseudo-tokens for efficient conditioning of a frozen LLM. Experiments on four public benchmarks show that MGSI substantially outperforms frozen-LLM baselines and remains competitive with strong multimodal methods. Further ablation and sensitivity analyses support the effectiveness of multi-granularity temporal modeling, text-guided refinement, and adaptive sentiment calibration.
CommentsAccepted to NLPCC 2026