arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低资源语言仇恨言论检测的大语言模型高效适配:以罗马乌尔都语为例的比较研究

Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu

Toneema Zubair, Muhammad Junaid Asif, Faisal Kamiran, Hafiz Hassan Saeed, Rana Fayyaz Ahmad

arXiv 2608.18142首次发表:更新:

AI 中文总结

本文针对罗马乌尔都语仇恨言论检测的低资源语言场景,对比评估Mistral等大语言模型,发现采用LoRA参数高效微调可显著提升性能,且计算效率优异。

AI 中文摘要

低资源语言(LRLs)的仇恨言论检测极具挑战性,原因在于缺乏标注数据、语言结构不规范以及标准化语法缺失,罗马乌尔都语就是典型案例——该语言被南亚民众广泛用于社交媒体,拼写变体多且缺乏语境一致的规范。本文旨在对罗马乌尔都语书写体系下仇恨言论检测(HSD)的大语言模型(LLMs)开展全面评估,并采用名为低秩适配(LoRA)的参数高效微调(PEFT)方法对这些模型进行微调。为评估零样本推理效果,本文将其与PEFT在不同Transformer模型(包括Mistral、LLaMA、Falcon及多语言BERT)上的表现进行基准测试。实验在包含超7.2万条标注评论的PURUTT(用于有毒评论与转写的乌尔都语-罗马乌尔都语平行语料库)数据集上开展。结果显示,零样本模型的表现处于中等水平(F1值为0.56),而仅更新模型的一小部分可训练参数就能显著提升分类性能(F1值大于0.93)。本文研究结果表明,PEFT兼具出色性能与优异计算效率,非常适用于低资源语言处理任务。

英文摘要

It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its language structure, and the lack of standardized grammar. A good example of such a challenge is Roman Urdu which is broadly used by South Asians on social media and has a high variation while lacking contextually consistent spellings. The objective of this paper is to conduct a comprehensive assessment of Large Language Models (LLMs) for Hate Speech Detection (HSD) in Roman Urdu script and fine-tune these models using the Parameter-Efficient Fine-Tuning (PEFT) method called Low-Rank Adaptation (LoRA). To evaluate zero-shot inference, we benchmarked it against PEFT on different transformer models, including Mistral, LLaMA, Falcon, and multilingual BERT. Experiments are conducted on the PURUTT (Parallel Urdu and Roman Urdu Corpus for Toxic Comments and Transliteration) dataset with over 72,000 annotated comments. The results suggest that zero shot models perform moderately (F1 = 0.56), but updating a small fraction of the model trainable parameters improves the classification performance significantly (F1 > 0.93). Our results have shown that PEFT delivers outstanding performance alongside excellent computational efficiency, making it highly suitable for low-resource language processing tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑