用于多模态情感解释的令牌-区域引导交叉注意力融合
Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation
浏览论文内容
中文总结 AI 辅助
针对孟加拉语模因政治意图检测难题,引入多模态交叉注意力融合框架,利用视觉语言模型提取文本,经跨模态多头注意力机制合成特征,还整合政治词汇表,实验表明该方法显著优于基线,Macro-F1达0.94,有效实现文本语义基于视觉证据。
中文摘要 AI 辅助
在数字时代,社交网络上多模态内容的自动分析对于理解公众情绪和信息传播至关重要。然而,对网络模因进行分类在计算上仍具有挑战性,尤其是对于孟加拉语等低资源语言。本文通过引入多模态交叉注意力融合框架来解决孟加拉语模因中政治意图的检测问题。首先利用视觉语言模型从有噪声的模因图像中提取高保真OCR文本,然后编码视觉和文本特征并通过跨模态多头注意力机制合成,还研究了特定领域政治词汇表的整合。在PoliMemeDecode1数据集上的实验表明,基于注意力的融合显著优于单模态基线和标准拼接方法,Macro-F1约为0.94,可有效将文本语义基于视觉证据,解释性分析进一步证实了这一点。
英文摘要
Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the intricate interplay between visual cues and embedded, often stylized, text, particularly in low-resource languages like Bengali Language. This paper addresses the detection of political intent in Bengali memes by introducing Multimodal Cross-Attention Fusion framework. We first leverage a Vision-Language Model to extract high-fidelity OCR text from noisy meme images. Subsequently, we encode visual and textual features and synthesize them through a cross-modal multi-head attention mechanism that aligns semantic tokens with visual regions. We also investigate the integration of a domain-specific political lexicon as a knowledge prior. Experimental evaluation on the PoliMemeDecode1 dataset shows that our attention-based fusion significantly outperforms unimodal baselines and standard concatenation methods, achieving a state-of-the-art Macro-F1 of approximately 0.94. Interpretability analyzes further confirm that the model effectively learns to ground textual semantics in visual evidence.