arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26143cs.CLcs.AI

超越准确率:对用于表情包仇恨言论检测的视觉-语言模型的定性分析

Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes

  • Islamic University of Technology(伊斯兰科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Muhammad Jawad Chowdhury, Adiba Hasan, Ishrak Hossain, Shahriar Ivan, Sabbir Ahmed

AI总结:

本研究通过定性分析LLaVA-7B等四个VLMs在零样本、少样本设置下的表现,揭示其检测仇恨表情包时忽略关键语境线索的问题,为优化模型提供了更深入的理解。

AI中文摘要:

表情包已成为个人分享关于当代社会与政治问题观点的有力工具,其匿名性与传播性使其成为传播仇恨的强大媒介,识别这类复杂且依赖语境的仇恨言论极为困难。尽管视觉-语言模型(VLMs)在多模态任务中表现出色,但它们往往会忽略语境、反讽及其他对识别仇恨表情包至关重要的微妙线索。本研究对四个最先进的VLMs——LLaVA-7B、Qwen-VL、GPT-4o mini和Claude 3 Haiku进行了定性分析,在零样本与少样本提示设置下评估这些模型,以探究语境框架如何影响其输出。本分析超越了简单的分类准确率,聚焦于对模型生成的理由进行定性评估,从而更深入地理解它们处理仇恨表情包时的思维过程与局限性。

英文摘要:

Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for spreading hate. It remains very difficult to identify such complex and context-dependent hate speech. Although they display excellent performance on multimodal tasks, vision-language models (VLMs) tend to ignore context, irony, and other subtle cues that play a key role in identifying hateful memes. In this work, we present a qualitative analysis of four state-of-the-art VLMs: LLaVA-7B, Qwen-VL, GPT-4o mini, and Claude 3 Haiku. We evaluate these models under zero-shot and few-shot prompting to examine how contextual framing influences their outputs. Our analysis goes beyond simple classification accuracy and focuses on a qualitative evaluation of the models' generated justifications, providing a more in-depth understanding of their thought processes and constraints when dealing with hateful memes.

↑