发表机构
INAOE; The University of Texas at El Paso(国家天体物理、光学与电子学研究所; 德克萨斯大学埃尔帕索分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MexHat数据集,包含约1000个墨西哥西班牙语视频片段,用于仇恨言论检测,涵盖三分类及细粒度子类别标注,并报告基线结果以揭示任务挑战。
AI 中文摘要
通过内容监控确保在线安全,已将仇恨言论检测提升为一项亟待解决的关键任务。本质上,该任务要求捕捉上下文线索,这对于准确理解内容意图至关重要。尽管针对该任务的自动化检测方法已取得显著进展,但非英语资源的稀缺性依然存在,限制了模型适应多模态内容中微妙、依赖上下文及与文化相关特性的能力。在本文中,我们介绍了MexHat,一个旨在捕捉墨西哥西班牙语语境下仇恨言论检测任务中语言和文化线索的视频数据集。我们的数据集包含约1000个视频片段,标注涵盖两项任务:一项三分类评估(无负面内容、冒犯性内容和仇恨言论内容),以及一项细粒度分类评估,包含三个仇恨言论子类别。数据集统计数据和基线结果凸显了该任务固有的挑战。免责声明:本文包含可能令部分读者感到不适的敏感内容。
英文摘要
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.
CommentsPreprint submitted to CIARP 2026