AI 中文总结
本研究针对多模态恶意梗图检测中存在的目标识别错误问题,提出HarmTrace框架,结合实体感知微调与CTPO方法,在Qwen3-VL-8B上JRA提升至52.51%,还构建了Meme3W数据集与JRA指标。
AI 中文摘要
多模态恶意梗图检测通常被建模为图像-文本恶意性分类任务,模型可能在正确预测恶意性的同时,错误识别被攻击的目标或其支撑证据。因此,我们将恶意梗图检测扩展为细粒度目标识别任务,需明确被攻击的目标类型、具体目标对象以及目标在梗图中的出现位置。模型需预测每个梗图的恶意性,对于恶意梗图,需输出目标类别、目标实体、文本提及内容和视觉区域。为支撑该任务,我们引入Meme3W数据集,该数据集整合了多个公开恶意梗图数据集,并为恶意实例提供经人工验证的标注。我们进一步提出联合记录准确率(Joint Record Accuracy, JRA),这是一种严格的记录级指标,要求恶意性标签与所有目标识别字段需联合正确。对代表性多模态大语言模型的实验表明,恶意性准确率与JRA之间存在显著差距。为缩小该差距,我们提出HarmTrace,一种锚定校准解耦优化框架。HarmTrace通过实体感知的监督微调强化目标实体监督,随后应用条件目标识别策略优化(Conditional Target-identification Policy Optimization, CTPO)来解耦恶意性与目标识别优势,将目标识别优化限制在恶意示例的标签正确响应范围内。CTPO使用虚拟正锚(Virtual Positive Anchor, VPA)作为目标识别优势归一化的完全正确参考。HarmTrace在所有评估的骨干模型上均提升了JRA和恶意性准确率,其中Qwen3-VL-8B骨干模型的JRA从17.58%提升至52.51%。我们的代码可在指定URL公开获取。
英文摘要
Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfulness while misidentifying the attacked target or its supporting evidence. We therefore extend harmful meme detection with fine-grained target identification, asking what type of target is attacked, who is targeted, and where the target appears in the meme. The model predicts harmfulness for every meme and, for harmful memes, outputs the target category, target entity, textual mention, and visual region. To support this task, we introduce Meme3W, which unifies multiple public harmful meme datasets and provides human-verified annotations for harmful instances. We further introduce Joint Record Accuracy (JRA), a strict record-level metric requiring the harmfulness label and all target-identification fields to be jointly correct. Experiments with representative multimodal large language models reveal a substantial gap between harmfulness accuracy and JRA. To narrow this gap, we propose HarmTrace, an anchor-calibrated decoupled optimization framework. HarmTrace strengthens target-entity supervision through entity-aware supervised fine-tuning. It then applies Conditional Target-identification Policy Optimization (CTPO) to decouple harmfulness and target-identification advantages, restricting target-identification optimization to label-correct responses for harmful examples. CTPO uses a Virtual Positive Anchor (VPA) as a fully correct reference for target-identification advantage normalization. HarmTrace improves both JRA and harmfulness accuracy across the evaluated backbones, with JRA on the Qwen3-VL-8B backbone increasing from 17.58\% to 52.51\%. Our code is publicly available at https://github.com/llly1234/HarmTrace-for-Harmful-Memes.