arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向社交媒体多模态情感分析的细粒度视觉预处理与双流时序建模

Fine-Grained Visual Preprocessing and Dual-Stream Temporal Modeling for Multimodal Sentiment Analysis on Social Media

Su Li, Yigong Zhang, Lei Xiong, Chune Li

arXiv 2609.07010首次发表:更新:

发表机构

School of Information Engineering, Kunming University(昆明大学信息工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态情感分析中视觉噪声和时序建模不足的问题,提出NAPS预处理、DS-TANet双流时序网络及DS-TAFNet融合模型,在CH-SIMS v2.0S上显著提升性能,证明改进视觉质量优于增加融合复杂度。

AI 中文摘要

多模态情感分析常因原始视频噪声和时序建模不足而呈现文本主导现象。本研究基于CH-SIMS v2.0S数据集,提出三项改进:NAPS流水线——一个七阶段系统,整合人脸跟踪、身份嵌入和归一化唇部运动分析以减少视觉噪声;DS-TANet,结合EfficientNetB2静态流、RAFT光流运动流、运动引导注意力及Bi-GRU时序建模;以及DS-TAFNet,通过拼接融合视觉和MacBERT-Base文本表示。采用NAPS后,静态视觉基线达到80.98%的Macro F1,与文本基线的80.55%相当;DS-TANet将视觉Macro F1提升至82.58%;DS-TAFNet达到87.49%的准确率和87.48%的Macro F1。这些结果表明,在数据有限的条件下,提升视觉输入质量和时序表示比增加融合复杂度更为有效。

英文摘要

Multimodal sentiment analysis often remains text-dominant due to raw-video noise and insufficient temporal modeling. Using CH-SIMS v2.0S, this study proposes three improvements: the NAPS pipeline---a seven-stage system integrating face tracking,identity embedding, and normalized lip-motion analysis to reduce visual noise;DS-TANet, combining an EfficientNetB2 static stream, RAFT optical-flow motion stream, motion-guided attention, and Bi-GRU temporal modeling; and DS-TAFNet, fusing visual and MacBERT-Base textual representations via concatenation fusion. With NAPS, the static visual baseline achieves 80.98\% Macro F1, comparable to the text baseline of 80.55\%; DS-TANet improves visual Macro F1 to 82.58\%;and DS-TAFNet achieves 87.49\% accuracy and 87.48\% Macro F1. These results demonstrate that improving visual input quality and temporal representation is more effective than increasing fusion complexity under limited-data conditions.

Comments35 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑