arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FAD-SA-GRU:通过特征增强自注意力GRU网络增强阿尔及利亚方言中的仇恨言论检测

FAD-SA-GRU: Enhancing Hate Speech Detection in Algerian Dialect Through Feature-Augmented Self-Attention GRU Networks

Sara Yakoubi, Ikram Khalfallah, Kenza Khelkhal, Dihia Lanasri

arXiv 2607.11279首次发表:更新:

发表机构

USTHB; ATM Mobilis Algiers(阿尔及尔科学技术大学; 阿尔及尔移动ATM公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究阿尔及利亚方言社交媒体上的仇恨言论检测,提出FAD-SA-GRU混合架构,通过多嵌入融合结合多种语义表示,经自注意力增强GRU编码器检测。实验表明该方法在多指标上优于基线,有效提升低资源方言仇恨言论检测效果。

AI 中文摘要

社交媒体平台的广泛应用改变了在线交流方式,但也助长了辱骂和仇恨内容的传播,仇恨言论检测成为自然语言处理的重要研究课题。本文研究阿尔及利亚阿拉伯方言(达里贾语)在社交媒体上的仇恨言论自动检测。由于该方言语言多样,任务颇具挑战。文中比较了四类文本分类方法,包括传统机器学习模型、基于循环神经网络的深度学习模型、基于Transformer的语言模型以及新型混合架构FAD-SA-GRU。实验表明,FAD-SA-GRU在阿尔及利亚达里贾语社交媒体评论的二元仇恨言论分类标注数据集上,优于所有基线模型,各项指标表现出色,证明了结合互补嵌入表示与基于注意力的序列建模对低资源方言阿拉伯语中仇恨言论检测的有效性。

英文摘要

The widespread adoption of social media platforms has transformed online communication by enabling users to exchange information and opinions instantly. However, these platforms have also facilitated the dissemination of abusive and hateful content, posing major social, psychological, and ethical challenges. Hate speech can incite discrimination, harassment, and violence against individuals or communities based on attributes such as ethnicity, religion, gender, nationality, or political affiliation. Consequently, automatic hate speech detection has become a major research topic in natural language processing (NLP) and an essential component of content moderation systems. This paper investigates automatic hate speech detection in the Algerian Arabic dialect (Darija) on social media. This task remains challenging because of the dialect's linguistic diversity, characterized by the coexistence of Arabic, French, and Arabizi (Arabic written using the Latin alphabet). We compare four categories of text classification approaches: (1) traditional machine learning models using TF-IDF features, (2) deep learning models based on recurrent neural networks, (3) Transformer-based language models, including DziriBERT and multilingual BERT, and (4) a novel hybrid architecture, FAD-SA-GRU, which combines semantic representations from DZ FastText, DZ AraVec, and DziriBERT through multi-embedding fusion, followed by a self-attention-enhanced GRU encoder. Experiments on an annotated dataset of Algerian Darija social media comments for binary hate speech classification show that FAD-SA-GRU outperforms all baselines, achieving 93.2% accuracy, 93.4% precision, 91.0% recall, 92.1% F1-score, and 97.0% ROC-AUC. Results demonstrate the effectiveness of combining complementary embedding representations with attention-based sequence modeling for robust hate speech detection in low-resource dialectal Arabic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑