arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先学后判:面向可解释仇恨梗检测的渐进式知识到决策对齐

Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection

Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia

arXiv 2609.19778首次发表:更新:

发表机构

Dalian University of Technology; RMIT University(大连理工大学; 皇家墨尔本理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有可解释仇恨梗检测中解释与预测耦合导致性能下降的问题,提出ProKDA方法,通过智能体知识构建与三阶段渐进对齐,在三个基准上取得最优检测性能并提供可解释决策。

AI 中文摘要

仇恨梗通过图像与文本之间的隐式交互传播辱骂性内容,对在线社区的安全构成严重威胁。近年来,多模态大语言模型已被广泛用于仇恨梗检测,并越来越多地被用于生成可解释的检测结果。然而,我们发现现有的“先解释后检测”方法往往在相同的训练过程中耦合解释生成与标签预测。这种耦合导致任务目标之间的干扰,使得检测性能受限,甚至比简单的SFT基线更差。为应对这些挑战,我们提出ProKDA,一种用于可解释仇恨梗检测的渐进式知识到决策对齐方法。受人类标注训练过程的启发,ProKDA首先使用智能体背景知识构建流水线来获取与梗理解相关的外部知识。随后,它采用三阶段训练策略,依次进行背景知识学习、仇恨性检测学习和仇恨性边界对齐。与先前联合优化两个任务的“先解释后检测”方法不同,ProKDA在每个阶段专注于单一训练目标。这种设计减少了两个任务之间的干扰,并逐步将背景知识转化为稳健的检测决策。在三个公开的仇恨梗基准上的实验表明,ProKDA取得了最先进的检测性能,并为仇恨梗审核提供了准确、可解释且有证据支持的决策。项目页面:此https URL。

英文摘要

Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explainable detection results. However, we find that existing explain-then-detect methods often couple explanation generation and label prediction within the same training process. This coupling causes interference between task objectives, leading to limited detection performance and even worse results than simple SFT baselines. To address these challenges, we propose ProKDA, a progressive knowledge-to-decision alignment method for explainable hateful meme detection. Inspired by the human annotation training process, ProKDA first uses an agentic background knowledge construction pipeline to obtain external knowledge related to meme understanding. It then adopts a three-stage training strategy that sequentially performs background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment. Unlike prior explain-then-detect methods that jointly optimize both tasks, ProKDA focuses on a single training objective at each stage. This design reduces interference between the two tasks and progressively transforms background knowledge into robust detection decisions. Experiments on three public hateful meme benchmarks show that ProKDA achieves state-of-the-art detection performance and provides accurate, explainable, and evidence-supported decisions for hateful meme moderation. Project page: https://meizhiyuan88666.github.io/prokda.

Comments26 pages, 16 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑