arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27844cs.CL

EvoHarmBench:通过类人迭代规避突破内容审核

EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

Ruijie Jian, Benlei Cui, Ting Ma, Haidong Ding, Kangwei Liu, Ziwen Xu, Longtao Huang, Hui Xue, Ziqiang Zhu, Junjie Li, Haiwen Hong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出首个动态对抗评估框架EvoHarmBench,针对5类违规229个语义子簇的5002个样本评估LLM审核模型,发现SOTA模型经12次迭代后规避成功率达80.3%,将推动内容安全评估转向动态模式。

中文摘要 AI 辅助

现有有害内容检测评估主要依赖静态基准,难以反映现实内容平台中用户会根据审核反馈不断修改表达的交互式对抗生态,这种不匹配导致离线基准分数与在线部署效果存在显著性能差距。据我们所知,我们提出了首个用于内容审核系统的动态对抗评估框架EvoHarmBench,该框架采用迭代优化循环,在语义簇层面演化规避策略,同时优化规避成功率与人类可读性。我们对现实审核系统中广泛使用的基于大语言模型(LLM)的防御模型进行了系统评估,评估涵盖从内容平台收集的5002个现实对抗样本中衍生的5个违规类别下的229个语义子簇。实验显示,即使是领先的商业系统也存在重大漏洞:在可读性约束下,经过12次优化迭代后,SOTA LLM审核器的攻击成功率达到80.3%。我们将发布完整的基准数据、评估框架及代码,以推动内容安全研究从静态基准评估转向动态对抗评估。

英文摘要

Existing evaluations of harmful content detection rely predominantly on static benchmarks, which struggle to reflect the interactive adversarial ecosystem of real-world content platforms where users continuously revise their expressions in response to moderation feedback. This mismatch creates a significant performance gap between offline benchmark scores and online deployment effectiveness. To the best of our knowledge, we present EvoHarmBench, the first dynamic adversarial evaluation framework for content moderation systems. The framework employs an iterative optimization loop that evolves evasion strategies at the semantic-cluster level, while simultaneously optimizing for evasion success and human readability. We systematically evaluate LLM-based defense models which are widely used in real world moderation systems. The evaluation covers 229 semantic sub-clusters across five violation categories, derived from 5,002 real-world adversarial samples collected from content platforms. Our experiments reveal substantial vulnerabilities even in leading commercial systems: after twelve optimization iterations, the attack success rate under readability constraints reaches 80.3% within SOTA LLM moderators. We will release the full benchmark data, evaluation framework, and code to encourage a shift from static benchmarking toward dynamic adversarial evaluation in content safety research.

发表机构

  • Yuvion Team, Alibaba Group(阿里巴巴集团Yuvion团队)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑