arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

洗白仇恨内容、抹黑无害内容:针对基于大语言模型的内容审核的标注者式反驳攻击

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

Junyu Lu, Kaiyuan Liu, Kaichun Wang, Jingyi Kang, Deyi Ji, Hailong Zhang, Lanyun Zhu, Qi Zhu, Bo Xu, Liang Yang, Hongfei Lin

arXiv 2608.22230首次发表:更新:

发表机构

Dalian University of Technology; Zhejiang University; University of Science and Technology of China; Tencent; Tongji University(大连理工大学; 浙江大学; 中国科学技术大学; 腾讯; 同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLMs内容审核提出标注者式反驳攻击,发现洗白仇恨内容与抹黑无害内容的攻击效果存在模型特异性不对称性,防御措施无法完全消除攻击影响,凸显需定向防御与反馈鲁棒性评估。

AI 中文摘要

大语言模型(LLMs)越来越多地被用于仇恨言论审核,通常在人机协同工作流程中,审核人员会在最终决策前提供反馈。这类反馈引入了两种操纵方向:将仇恨内容洗白为正常内容,以及将正常内容抹黑为仇恨内容。本研究探究了初始判断正确的模型对标注者式反驳的敏感性,并分析了攻击效果是否因操纵方向而异。我们提出了一种重判协议,该协议将直接矛盾扩展至决策边界扰动和对抗性理由。在两个仇恨言论数据集上对多个LLMs开展的实验表明,标注者式反驳会大幅降低审核性能,在多轮设置中效果更强。结果进一步揭示,在所有攻击配置下,洗白和抹黑之间存在稳定的、特定于模型的不对称性,表明存在不同的定向脆弱性模式。明确的推理提示和防御指令可降低这些影响,但无法将其消除。这些发现凸显了人机协同审核工作流程中对定向感知防御措施以及专用反馈鲁棒性评估的需求。

英文摘要

Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderation performance, with stronger effects in multi-turn settings. The results further reveal stable, model-specific asymmetries between whitewashing and smearing across attack configurations, indicating distinct directional vulnerability patterns. Explicit reasoning prompts and defensive instructions reduce these effects but do not eliminate them. These findings highlight the need for direction-aware safeguards and dedicated feedback-robustness evaluation in human--AI moderation workflows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑