arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16393cs.CL

ParsHate:波斯语仇恨与目标检测基准数据集

ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

Zahra Bokaei, Walid Magdy, Bonnie Webber

首次发表
浏览论文内容

中文总结 AI 辅助

ParsHate是首个覆盖2013-2022年波斯语推文的十年仇恨言论检测基准,含31%仇恨内容及多标签目标识别,现有SOTA模型性能中等(F1 79%),目标识别挑战大(宏F1 25.5%)。

中文摘要 AI 辅助

我们介绍了ParsHate,一个包含10,000条波斯语推文(时间跨度2013-2022年)的人工标注数据集,代表了波斯语仇恨言论检测领域首个长达十年的基准。该数据集包含31%的仇恨内容,支持仇恨检测以及跨七个结构化目标类别的多标签细粒度目标识别。ParsHate还区分显性仇恨和隐性仇恨,标记显性目标和隐性目标,并提供跨度级理由。数据收集结合了随机和时间分层抽样,以减少关键词驱动的偏差,同时保持自然标签分布。将波斯语仇恨言论检测的SOTA模型应用于ParsHate,显示出中等性能(79%的F1分数),尤其是在早期年份的样本上,而目标识别性能较低(25.5%的宏F1)。这强调了ParsHate中仇恨言论采样的多样性及其挑战性,需要更先进的方法以获得更好的性能。该数据集已公开提供。

英文摘要

We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identification across seven structured target categories. ParsHate also distinguishes explicit and implicit hate, marks explicit and implicit targets, and provides span-level rationales. Data collection combines random and score-stratified temporal sampling to reduce keyword-driven bias while preserving natural label distributions. Applying SOTA models for Persian hate-speech detection on ParsHate shows moderate performance (79% F1), especially with samples from earlier years, and low performance with target identification (25.5% macro-F1). This emphasizes the diverse sampling of hate speech in ParsHate and its challenging nature that requires more advanced methods for better performance. Dataset is made publicly available.

发表机构

  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑