arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemeTAG:通过标签嵌入重建实现关键词驱动的模因分类

MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction

Akshit Sharma, Prashant W. Patil

arXiv 2609.20962首次发表:更新:

发表机构

Indian Institute of Technology Guwahati(印度理工学院古瓦哈提分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MemeTAG提出双目标框架,利用视觉-语言模型生成关键词并经ATIN模块聚合为语义嵌入,通过辅助重建损失对齐视觉与文本特征,在多个数据集上超越现有方法。

AI 中文摘要

有害互联网模因的泛滥构成了重大的社会威胁,然而由于其内容具有微妙的多模态特性,其自动化分类仍然是一个艰巨的算法挑战。为解决这一问题,我们提出了MemeTAG,一个新颖的双目标框架,开创了关键词感知的模因分类方法。我们的核心创新是一个两部分组成的语义引导机制:首先,我们利用预训练的视觉-语言模型生成一组描述性关键词,以捕捉高层语义。其次,我们引入了聚合标签推理网络(ATIN),一个基于注意力的模块,将这些关键词提炼为单一的丰富语义嵌入。该嵌入作为新颖辅助重建损失的目标,迫使模型学习深度对齐的视觉和文本特征。这种方法结合高效的三阶段训练策略,在HarMeme、Hateful Memes Challenge(HMC)和PrideMM数据集上确立了新的最先进水平,明显优于现有的最先进方法。

英文摘要

The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers a keyword-aware approach to meme classification. Our core innovation is a two-part semantic guidance mechanism: first, we leverage a pretrained Vision-Language Model to generate a set of descriptive keywords, that capture the high-level semantics. Second, we introduce the Aggregated Tag Inference Network (ATIN), an attention-based module that distills these keywords into a single, rich semantic embedding. This embedding serves as a target for a novel auxiliary reconstruction loss, which compels the model to learn deeply aligned visual and textual features. This approach, combined with an efficient three-stage training strategy, establishes a new state-of-the-art on the HarMeme, Hateful Memes Challenge (HMC), and PrideMM datasets, decisively outperforming existing state-of-the-art methods.

Comments10 pages, 3 figures; published in the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

Journal refProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 7679-7688

DOI:10.1109/WACV61042.2026.00741

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑