超越基准:揭示孟加拉语仇恨言论检测中的隐藏危机
Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection
- Rajshahi University of Engineering and Technology(拉杰沙希工程技术大学)
- University of York(约克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究孟加拉语仇恨言论检测,训练六种架构于基准和多源数据集,在外部验证。发现基准训练模型泛化危机,如孟加拉语BERT在外部数据集性能下降,强调需开发适应性框架,为设计检测系统提供见解。
AI中文摘要:
仇恨言论在不同社交媒体平台上的传播对网络安全和道德审核构成重大担忧。自动检测仇恨言论仍是一项具有挑战性的任务,尤其是在孟加拉语这种资源匮乏的语言中,因为存在文化背景、隐含表达和非正式语言模式。本研究旨在通过诊断基准训练模型无法识别隐含的、依赖上下文的仇恨言论的方式和原因,揭示孟加拉语仇恨言论检测系统的危机。六种架构在基准数据集和合并的多源数据集上训练,然后在从脸书、推特和YouTube收集的注释数据集上进行外部验证。结果表明,孟加拉语BERT在基准数据集上F1分数为91.4%,但在外部数据集上降至75.3%,对于涉及讽刺和表情符号的隐含仇恨言论降至63.4%。表情符号感知预处理将隐含仇恨言论检测提高了12%,而去除表情符号导致性能显著下降。政治化或讽刺性评论中的频繁错误分类揭示了过度监管的风险。本研究不仅揭示了由于隐含、文化嵌入和充满表情符号的表达导致的泛化危机,还强调了开发适应性、表情符号感知和文化基础框架的必要性,这些框架在确保道德审核的同时保留表达自由。研究结果为研究人员、社交媒体平台和政策制定者设计更具上下文敏感性的低资源语言仇恨言论检测系统提供了见解。
英文摘要:
The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced languages like Bangla, due to cultural context, implicit expressions, and informal linguistic patterns. This study aimed to expose the crisis of Bangla HS detection systems by diagnosing how and why benchmark-trained models fail to identify implicit, context-dependent HS. Six architectures (FastText + CNN, FastText + LSTM, FastText + BiLSTM, BanglaBERT, BanglaBERT + CNN, and BanglaBERT + BiLSTM) were trained on benchmark datasets (about 75,000 posts) and a merged multi-source dataset (about 120,000 posts), then externally validated on an annotated dataset (about 200 posts) collected from Facebook, Twitter, and YouTube, labeled as HS and non-HS, where HS was further categorized as explicit and implicit. BanglaBERT achieved an F1-score of 91.4% on benchmark datasets but declined to 75.3% on the external set and 63.4% for implicit HS involving sarcasm and emojis. The accuracy of FastText + CNN dropped from 78.0% to 51.2% under similar conditions. Emoji-aware preprocessing improved implicit HS detection by up to 12%, whereas emoji removal caused a notable decline in performance (F1: 0.75 to 0.63). Frequent misclassifications in politically charged or satirical comments revealed over-policing risks. This study not only exposes the generalization crisis due to implicit, culturally embedded, and emoji-laden expressions but also underscores the need for developing adaptive, emoji-aware, and culturally grounded frameworks that ensure ethical moderation while preserving freedom of expression. Findings of this study provide insights for researchers, SMPs, and policymakers to design more context-sensitive HS detection systems for low-resource languages.