发表机构
Hasso Plattner Institute, University of Potsdam(波茨坦大学哈索·普拉特纳研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对HateBench开展可复现性研究,重建数据集生成流程并扩展基准,确认原始总体发现,发现新LLM有安全措施,Text-Moderation对抗仇恨运动鲁棒性略升。
AI 中文摘要
随着大语言模型(Large Language Models,LLM)降低了自动内容生成的门槛,其产生仇恨言论的潜力对数字安全构成了重大挑战。本文对Shen等人的HateBench论文开展了可复现性研究,调查通常在人工撰写数据上训练的现有仇恨言论检测器是否能泛化到LLM生成的仇恨内容,并评估其报告的弱点是否随时间保持稳定、是否对不断演变的组件具有鲁棒性。我们使用现代LLM独立重建了原始数据集生成流程,并将该基准扩展为包含最近发布的模型和更新后的检测器版本。我们在当前条件下的独立评估发现,针对较新的LLM,已部署了安全措施以防止有害内容的生成。我们还复现了两种复杂仇恨运动类型的结果。虽然原始发现似乎因数据集偏差而被略微高估,但总体发现可得到确认。最后,我们将Text-Moderation与较新的omni-Moderation进行比较,发现其对抗对抗性仇恨运动的鲁棒性略有提升。本研究通过明确哪些检测器漏洞持续存在,为社区提供了内容 moderation 测量持久性的相关信息。
英文摘要
As Large Language Models (LLMs) lower the barrier for au- tomated content generation, the potential for producing hate speech poses a significant challenge for digital safety. This paper presents a reproducibility study of the HateBench paper by Shen et al., investigating whether existing hate speech detectors, typically trained on human-authored data, generalize to LLM-generated hateful content, and evaluating whether their reported weaknesses are stable over time and robust to evolving components. We independently reconstruct the original dataset genera- tion pipeline using modern LLMs and extend the benchmark to include recently released models and updated detector versions. Our independent assessment under current con- ditions finds that for newer LLMs, safeguards have been put into place to prevent the generation of harmful content. We also replicate the results for two sophisticated types of hate campaigns. While the original findings seem to have been overestimated slightly due to bias in the datasets, the overall findings can be confirmed. Finally, we compare text- Moderation against the newer omni-Moderation and find that its robustness against adversarial hate campaigns has improved slightly. By clarifying which detector vulnerabil- ities persist, this study informs the community about the longevity of content moderation measurements.