arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉引导的文本提示调优用于多模态情感分析

Vision-Guided Text Prompt Tuning for Multimodal Sentiment Analysis

Xiaoran Kou, Jingyi Wu, Peng Sun, Yang Liu, Hong Chen

arXiv 2609.06497首次发表:更新:

发表机构

Tongji University; Fudan University; Duke Kunshan University(同济大学; 复旦大学; 昆山杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态情感分析中视觉噪声和微调成本高的问题,提出视觉引导的文本提示调优(VG-TPT),通过层间自适应提示将视觉线索注入冻结BERT,仅更新240万参数,在CMU-MOSEI和CMU-MOSI上取得竞争性性能。

AI 中文摘要

多模态情感分析需要对言语语义和非言语情感线索进行有效建模。一个核心挑战是以可控、自适应且参数高效的方式,用视觉面部证据来校准以文本为中心的情感理解。文本通常作为语义锚点,而视觉线索为模糊或隐含的表达提供补充证据;然而,不加区分的融合可能引入视觉噪声并扭曲文本语义。此外,完全微调大型视觉和文本编码器成本高昂,且容易在有限且依赖场景的MSA基准上过拟合。为解决这些问题,我们提出了视觉引导的文本提示调优(VG-TPT),将视觉-文本情感建模形式化为对冻结文本表示的可控视觉校准。VG-TPT通过逐层自适应提示将视觉情感线索注入冻结的BERT编码器,而非依赖后期特征融合或完全骨干网络调优。一个协同引导路由器根据不断演化的文本状态和视觉引导特征,从可训练的提示库中组合提示,实现样本特定和层特定的调制。在CMU-MOSEI和CMU-MOSI上的实验表明,VG-TPT持续优于仅文本基线,并与多种全模态方法相比达到竞争性或更优性能,同时仅更新240万可训练参数。代码可在该https URL获取。

英文摘要

Multimodal sentiment analysis requires effective modeling of both verbal semantics and non-verbal affective cues. A central challenge is to calibrate text-centered sentiment understanding with visual facial evidence in a controlled, adaptive, and parameter-efficient manner. Text usually serves as the semantic anchor, whereas visual cues provide complementary evidence for ambiguous or implicit expressions; however, indiscriminate fusion may introduce visual noise and distort textual semantics. Moreover, fully fine-tuning large visual and textual encoders is costly and prone to overfitting on limited and scenario-dependent MSA benchmarks. To address these issues, we propose Vision-Guided Text Prompt Tuning (VG-TPT), which formulates visual-text sentiment modeling as controllable visual calibration of frozen text representations. VG-TPT injects visual affective cues into a frozen BERT encoder through layer-wise adaptive prompts, rather than relying on late-stage feature fusion or full backbone tuning. A co-guided router composes prompts from a trainable prompt bank according to both the evolving text state and the visual guidance feature, enabling sample-specific and layer-specific modulation. Experiments on CMU-MOSEI and CMU-MOSI show that VG-TPT consistently improves over text-only baselines and achieves competitive or superior performance compared with several full-modality methods, while updating only 2.4M trainable parameters. The code is available at https://github.com/ma-tubu/VG-TPT.

CommentsThis paper has been accepted to IEEE MMSP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑