发表机构
Qatar Computing Research Institute; Hamad Bin Khalifa University; Northwestern University in Qatar(卡塔尔计算研究所; 哈马德·本·哈利法大学; 卡塔尔西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出首个带细粒度多标签标注的阿拉伯语仇恨表情包基准AHA-Memes,构建含5000张人工标注及6.6万张银标的数据集,对多类模型基准测试并发布资源以推动相关研究。
AI 中文摘要
仇恨表情包是一种日益增长的多模态网络危害形式,其敌意意图通常通过图像、文本、文化典故和隐性目标的联合解读来传达。尽管高资源语言的仇恨表情包检测已取得进展,但阿拉伯语领域的研究仍不充分,现有表情包资源主要聚焦于宣传内容或粗粒度有害内容标签。我们推出AHA-Memes(阿拉伯语仇恨表情包),据我们所知,这是首个具有细粒度多标签标注的大规模阿拉伯语仇恨表情包基准。该数据集包含5000张经人工标注的表情包,采用捕获仇恨类型(即攻击策略)的分类体系;我们还提供约6.6万张银标表情包以支持未来研究。我们对仅文本、仅图像、后期融合多模态模型,以及少样本上下文学习(ICL)、开放和闭源权重的视觉语言模型(VLMs)在零样本和微调设置下进行了基准测试。我们的结果建立了强基线,并凸显了基于文化的阿拉伯语仇恨表情包检测中的关键挑战。我们发布了该数据集、标注指南和评估脚本以支持未来研究。警告:本文包含可能令读者不适的示例。
英文摘要
Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels. We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label annotations. The dataset includes 5K manually annotated memes using a taxonomy that captures hate types, i.e., attack strategies. We further provide ~66K silver-labeled memes to support future studies. We benchmark text-only, image-only, and late-fusion multimodal models, as well as few-shot in-context learning (ICL) and open- and closed-weight Vision-Language Models (VLMs) under zero-shot and fine-tuning settings. Our results establish strong baselines and highlight key challenges in culturally grounded Arabic hateful meme detection. We release the dataset, annotation guidelines, and evaluation scripts to support future research. WARNING: This paper contains examples that may be disturbing to readers.
Comments26 pages, 14 figures, 15 tables