arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MVFA:一种用于情感分析和情绪识别的多视图文本引导多模态融合LLM适配器

MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition

Pengfei Shao, Jisheng Dang, Jiawen Fang, Ning Liu, Wencan Zhang, Bimei Wang, Jingwen Zhao, Jianhuang Lai, Qi Tian, Tat-Seng Chua

arXiv 2609.06188首次发表:更新:

发表机构

Guangdong University of Foreign Studies; Lanzhou University; Sun Yat-sen University; Huawei; National University of Singapore(广东外语外贸大学; 兰州大学; 中山大学; 华为; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MVFA,一种参数高效的多视图文本引导多模态融合适配器,通过多种池化构建文本视图并引导跨模态融合,在三个数据集上取得最先进性能。

AI 中文摘要

对话中的多模态情感分析和情绪识别需要有效建模文本、声学和视觉模态之间的异构交互。尽管大型语言模型(LLM)提供了强大的语言理解能力,但将其适配到多模态情感计算仍然具有挑战性:全模型微调在计算上代价高昂,而许多现有的轻量级适配器在跨模态融合过程中未能保留丰富的文本线索。为了解决这些局限性,我们提出了多视图文本引导的多模态融合适配器(MVFA),这是一种参数高效的框架,能够增强冻结的LLM的强大多模态推理能力。MVFA首先通过最大池化、平均池化和注意力池化构建互补的文本视图;这些视图随后引导与音频和视觉特征的跨模态交互。融合后的多模态表示随后通过增强的Q-Former融合模块压缩为一组紧凑的可学习伪标记。以ChatGLM3-6B-base作为主要骨干网络,我们进一步在LLaMA2-7B和Qwen3-8B上验证MVFA,以检验其在多个冻结LLM骨干网络上的可移植性。MVFA在三个具有挑战性的数据集上进行了评估:CH-SIMS V2.0、MELD和CHERMA。实验结果表明,MVFA在仅更新一小部分参数的情况下,在关键指标上达到了最先进的性能。具体而言,它在CH-SIMS V2.0上达到了84.62%的Acc2和84.59%的F1,在MELD上达到了67.36%的Acc和66.03%的WF1,在CHERMA上达到了74.66%的Acc。这些发现确立了多视图文本引导融合作为情感计算中参数高效多模态LLM适配的一种有效且可扩展的范式。代码可在以下网址公开获取:此https URL。

英文摘要

Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic, and visual modalities. Although large language models (LLMs) offer powerful language understanding, adapting them to multimodal affective computing remains challenging: full-model fine-tuning is computationally prohibitive, while many existing lightweight adapters fail to preserve rich textual cues during cross-modal fusion. To address these limitations, we propose the multi-view text-guided multimodal fusion adapter (MVFA), a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability. MVFA first constructs complementary text views via max pooling, mean pooling, and attention pooling; these views then guide cross-modal interactions with audio and visual features. The fused multimodal representations are subsequently compressed into a compact set of learnable pseudo-tokens through an Enhanced Q-Former Fusion Module. Using ChatGLM3-6B-base as the primary backbone, we further validate MVFA on LLaMA2-7B and Qwen3-8B to examine its portability across multiple frozen LLM backbones. MVFA is evaluated on three challenging datasets: CH-SIMS V2.0, MELD, and CHERMA. Experimental results demonstrate that MVFA achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters. Specifically, it attains 84.62\% Acc2 and 84.59\% F1 on CH-SIMS V2.0, 67.36\% Acc and 66.03\% WF1 on MELD, and 74.66\% Acc on CHERMA. These findings establish multi-view text-guided fusion as an effective and scalable paradigm for parameter-efficient multimodal LLM adaptation in affective computing. The code is publicly available at https://github.com/Overwhelm1208/MVFA.

Comments10 pages, 4 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑