发表机构
Indian Institute of Technology Guwahati(印度理工学院古瓦哈提分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MemeTAG提出双目标框架,利用视觉-语言模型生成关键词并经ATIN模块聚合为语义嵌入,通过辅助重建损失对齐视觉与文本特征,在多个数据集上超越现有方法。
AI 中文摘要
有害互联网模因的泛滥构成了重大的社会威胁,然而由于其内容具有微妙的多模态特性,其自动化分类仍然是一个艰巨的算法挑战。为解决这一问题,我们提出了MemeTAG,一个新颖的双目标框架,开创了关键词感知的模因分类方法。我们的核心创新是一个两部分组成的语义引导机制:首先,我们利用预训练的视觉-语言模型生成一组描述性关键词,以捕捉高层语义。其次,我们引入了聚合标签推理网络(ATIN),一个基于注意力的模块,将这些关键词提炼为单一的丰富语义嵌入。该嵌入作为新颖辅助重建损失的目标,迫使模型学习深度对齐的视觉和文本特征。这种方法结合高效的三阶段训练策略,在HarMeme、Hateful Memes Challenge(HMC)和PrideMM数据集上确立了新的最先进水平,明显优于现有的最先进方法。
英文摘要
The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers a keyword-aware approach to meme classification. Our core innovation is a two-part semantic guidance mechanism: first, we leverage a pretrained Vision-Language Model to generate a set of descriptive keywords, that capture the high-level semantics. Second, we introduce the Aggregated Tag Inference Network (ATIN), an attention-based module that distills these keywords into a single, rich semantic embedding. This embedding serves as a target for a novel auxiliary reconstruction loss, which compels the model to learn deeply aligned visual and textual features. This approach, combined with an efficient three-stage training strategy, establishes a new state-of-the-art on the HarMeme, Hateful Memes Challenge (HMC), and PrideMM datasets, decisively outperforming existing state-of-the-art methods.
Comments10 pages, 3 figures; published in the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026
Journal refProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 7679-7688
DOI:10.1109/WACV61042.2026.00741