arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多语言仇恨言论检测的训练时可解释性:对齐模型推理与人类理由

Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

Muhammad Deedahwar Mazhar Qureshi, Sannaan Khan, Muhammad Atif Qureshi, Wael Rashwan

arXiv 2608.26125首次发表:更新:

发表机构

Technological University Dublin; National University of Sciences and Technology; Maynooth University(都柏林理工大学; 国家科技大学; 梅努斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多语言仇恨言论检测的不透明问题,提出训练时可解释性框架,对齐模型推理与人类理由,经HateXplain、BullySent数据集及多方法评估,可提升分类性能与解释质量,助力多语言文化敏感的内容审核。

AI 中文摘要

针对穆斯林群体的网络仇恨言论常以具有文化编码的多语言形式出现,可规避传统AI审核系统。这类系统虽具备准确性,但仍不透明,且存在偏见、过度审查或审核不足的风险,尤其在脱离社会文化背景时问题更突出。我们提出一种训练时可解释性框架,将模型推理与人工标注的理由对齐,同时提升分类性能与可解释性。在HateXplain(英语)和BullySent(印地语-英语混合语)数据集上对方法进行评估,这两个数据集反映了反穆斯林仇恨在两种语言中的普遍性。使用LIME、Integrated Gradients、Grad X Input和注意力机制,我们评估了准确性、解释质量及跨方法一致性。结果表明,基于梯度和注意力的正则化可提升F值,增强合理性与忠实度,并能捕捉文化特定线索以检测隐含的反穆斯林仇恨,为实现多语言、具文化意识的内容审核提供了路径。

英文摘要

Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability. Our approach is evaluated on HateXplain (English) and BullySent (Hinglish), reflecting the prevalence of anti-Muslim hate across both languages. Using LIME, Integrated Gradients, Grad X Input, and attention, we assess accuracy, explanation quality, and cross-method agreement. Results show that gradient- and attention-based regularization improve F-scores, enhance plausibility and faithfulness, and capture culturally specific cues for detecting implicit anti-Muslim hate, offering a path toward multilingual, culturally aware content moderation.

CommentsAccepted at NeurIPS Workshops 2025

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑